{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Building CNN From Scratch With Keras For Cancer Detection\n***By Kris Smith***\n\n* [Link to Project Github](https://github.com/kristopher-smith/CNN-Cancer-Detection/tree/main)\n\n* Great resource to help with visualizing CNN's: [Visualizing and Understanding Convolutional Networks](https://arxiv.org/pdf/1311.2901.pdf)","metadata":{}},{"cell_type":"markdown","source":"---\n# Problem Description\n\n***For this competition we must create an algorithm to identify metastatic cancer in small image patches taken from larger digital pathology scans. The data for this competition is a slightly modified version of the PatchCamelyon (PCam) benchmark dataset (the original PCam dataset contains duplicate images due to its probabilistic sampling, however, the version presented on Kaggle does not contain duplicates).***\n\n***This is an important problem because metastatic cells are an indicator of stage IV cancer. When cancer has \"metasticised\" it has left the original tumor/growth and traveled elsewhere throughout the body. This means it is harder to treat and the probability of fatality increases significantly. Identifying metastatic cells on tissue other than the original cancer site therefore helps to confirm this diagnosis allowing more effective and timely treatments/prognosis.***","metadata":{}},{"cell_type":"markdown","source":"## Import Libraries","metadata":{}},{"cell_type":"code","source":"import numpy as np, pandas as pd, tensorflow as tf\nimport os\nimport cv2\nfrom PIL import Image \nfrom glob import glob \nfrom keras.callbacks import *\nfrom keras.preprocessing.image import ImageDataGenerator\nfrom keras.models import *\nfrom tensorflow.keras.models import load_model\nfrom keras.layers.normalization import *\nfrom keras.layers.convolutional import *\nfrom keras.layers.core import *\nimport matplotlib.pyplot as plt\nfrom matplotlib.pyplot import imread\n%matplotlib inline\n\n\nprint(os.listdir(\"../input/histopathologic-cancer-detection\"))\n# tf.debugging.set_log_device_placement(True)\nprint(\"Num GPUs Available: \", len(tf.config.list_physical_devices('GPU')))\n#############################################################","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2023-05-29T21:48:23.204735Z","iopub.execute_input":"2023-05-29T21:48:23.205106Z","iopub.status.idle":"2023-05-29T21:48:26.060358Z","shell.execute_reply.started":"2023-05-29T21:48:23.205073Z","shell.execute_reply":"2023-05-29T21:48:26.059365Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"---\n\n# Data","metadata":{}},{"cell_type":"markdown","source":"## Data Description\n\n***In the authors words:***\n\n_\"[PCam] packs the clinically-relevant task of metastasis detection into a straight-forward binary image classification task, akin to CIFAR-10 and MNIST. Models can easily be trained on a single GPU in a couple hours, and achieve competitive scores in the Camelyon16 tasks of tumor detection and whole-slide image diagnosis. Furthermore, the balance between task-difficulty and tractability makes it a prime suspect for fundamental machine learning research on topics as active learning, model uncertainty, and explainability.\"_\n\n***Submissions are evaluated on area under the ROC curve between the predicted probability and the observed target.***","metadata":{}},{"cell_type":"markdown","source":"## EDA","metadata":{}},{"cell_type":"markdown","source":"### Load Data","metadata":{}},{"cell_type":"code","source":"print(f'Total Number of Samples in Training Data = {len(os.listdir(\"../input/histopathologic-cancer-detection/train\"))}')\nprint(f'Total Number of Samples in Test Data = {len(os.listdir(\"../input/histopathologic-cancer-detection/test\"))}')","metadata":{"execution":{"iopub.status.busy":"2023-05-29T21:48:26.066955Z","iopub.execute_input":"2023-05-29T21:48:26.067725Z","iopub.status.idle":"2023-05-29T21:48:26.277448Z","shell.execute_reply.started":"2023-05-29T21:48:26.067685Z","shell.execute_reply":"2023-05-29T21:48:26.276326Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"***Since the labels and images are provided to us as seperate files we need to read them both into a dataframe for processing and training easily.***\n\n***The below cell reads in the image and label data and concatenates the filepath including each images unique id with its image filetype and label.***","metadata":{}},{"cell_type":"code","source":"path = \"../input/histopathologic-cancer-detection/\"\ntrain_path = path + 'train/'\ntest_path = path + 'test/'\n\n### Read data into pandas frames ###\n## Load training and testing file names into frames\ndf = pd.DataFrame({'path': glob(os.path.join(train_path,'*.tif'))}) \ndf_test = pd.DataFrame({'path': glob(os.path.join(test_path,'*.tif'))}) \n\n## Isolate the id from file names and store as id features in seperate column\ndf['id'] = df.path.map(lambda x: x.split('/')[4].split(\".\")[0]) \n\n## Read in training labels\nlabels = pd.read_csv('../input/histopathologic-cancer-detection/train_labels.csv')\n\n## Add Labels to training data \ndf = df.merge(labels, on = \"id\") # merge labels and filepaths\ndf.sample(7)","metadata":{"execution":{"iopub.status.busy":"2023-05-29T21:48:26.279289Z","iopub.execute_input":"2023-05-29T21:48:26.279675Z","iopub.status.idle":"2023-05-29T21:48:28.022095Z","shell.execute_reply.started":"2023-05-29T21:48:26.279638Z","shell.execute_reply":"2023-05-29T21:48:28.020913Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Class Distributions","metadata":{}},{"cell_type":"code","source":"labels['label'].value_counts().plot(kind='barh', title='Class Distribution in Training Data')","metadata":{"jupyter":{"source_hidden":true},"execution":{"iopub.status.busy":"2023-05-29T21:48:28.025429Z","iopub.execute_input":"2023-05-29T21:48:28.026478Z","iopub.status.idle":"2023-05-29T21:48:28.252947Z","shell.execute_reply.started":"2023-05-29T21:48:28.026443Z","shell.execute_reply":"2023-05-29T21:48:28.251950Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"n = len(labels)\nnum_pos = labels['label'].sum()\nnum_neg = n - num_pos\nprop_pos = round(num_pos/n, 2)\nprop_neg = round(num_neg/n, 2)\n\nprint(f'Total Number of Positive Samples in Training Data = {num_pos} ')\nprint(f'Total Number of Negative Samples in Training Data = {num_neg} ')\nprint(f'Proportion of Positive Samples in Training Data = {prop_pos}%')\nprint(f'Proportion of Negative Samples in Training Data = {prop_neg}%')","metadata":{"execution":{"iopub.status.busy":"2023-05-29T21:48:28.254729Z","iopub.execute_input":"2023-05-29T21:48:28.256718Z","iopub.status.idle":"2023-05-29T21:48:28.264687Z","shell.execute_reply.started":"2023-05-29T21:48:28.256678Z","shell.execute_reply":"2023-05-29T21:48:28.263741Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"***Looks like there is a class imbalance but we seem to have a decent amount of data for each class so I decided to proceed without adjusting the data and see how the model works with this imbalanced data.***","metadata":{}},{"cell_type":"markdown","source":"## Inspect Some Samples From Both Positive and Negative Classes","metadata":{}},{"cell_type":"code","source":"## Filter dataframe to get a subset of each class\nclass0 = df[df['label'] == 0]\nclass1 = df[df['label'] == 1]\n\nfor image in range(5):\n    ## Select a random sample from each class\n    sample0 = class0.sample(1)\n    sample1 = class1.sample(1)\n\n    ## Open image files and plot\n    img_path0 = sample0['path'].values[0]\n    img_path1 = sample1['path'].values[0]\n\n    img0 = Image.open(img_path0)\n    img1 = Image.open(img_path1)\n\n    fig, ax = plt.subplots(1, 2, figsize=(5, 2))\n\n    ax[0].imshow(img0)\n    ax[0].set_title('Class 0')\n    ax[0].axis('off')\n\n    ax[1].imshow(img1)\n    ax[1].set_title('Class 1')\n    ax[1].axis('off')\n\n    plt.show()","metadata":{"execution":{"iopub.status.busy":"2023-05-29T21:48:28.266109Z","iopub.execute_input":"2023-05-29T21:48:28.266921Z","iopub.status.idle":"2023-05-29T21:48:29.361837Z","shell.execute_reply.started":"2023-05-29T21:48:28.266886Z","shell.execute_reply":"2023-05-29T21:48:29.360939Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"***According to the five minutes of extensive research I did on the interwebs, apparently metastatic cells should appear a different size, shape, or pattern compared to the healthy cells around them.***\n\n***Inspecting these images I am realizing that I am not a medical professional and therefore I can see no consistent differences between the positive and negative class images. Therefore I will be relying on the labels and the model to distinguish between the two.***","metadata":{}},{"cell_type":"markdown","source":"---\n# Feature Engineering","metadata":{}},{"cell_type":"markdown","source":"## Pre-Processing\n\n***I have chosen to split the training data into 90/10% training/validation sets respectively.***\n\n***For creating more out of the training data I have augmented the images with horizontal and vertical flips, as well as rotating within a certain range. This will give the model more to train on and since the images are direction ignorant, manipulating them in this manner preserves the data.***","metadata":{}},{"cell_type":"code","source":"df = pd.read_csv('../input/histopathologic-cancer-detection/train_labels.csv')\n# df = df.sample(12800*10) ## Smaller set for quick experimentation\n\n## ImageDataGenerator method requires string data type labels\ndf['label'] = df['label'].astype(str)\n\ndef append_ext(ID):\n    return(ID+\".tif\")\n\n\ndf[\"id\"] = df[\"id\"].apply(append_ext)\n\n\n\ntrain_datagen = ImageDataGenerator(\n#         horizontal_flip=True,\n#         vertical_flip=True,                               \n#         rotation_range=15,\n        rescale=1./255, ## Converts RGB values from range within [0, 255] ===> range within [0, 1]\n        validation_split=0.10\n    \n)\n\ntest_datagen = ImageDataGenerator(rescale=1./255)\n\ntrain_path = '../input/histopathologic-cancer-detection/train'\nvalid_path = '../input/histopathologic-cancer-detection/train'\n\ntrain_generator = train_datagen.flow_from_dataframe(\n                dataframe=df,\n                directory=train_path,\n                x_col = 'id',\n                y_col = 'label',\n                subset='training',\n                target_size=(96, 96),\n                batch_size=128,\n                class_mode='binary'\n                )\n\nvalidation_generator = train_datagen.flow_from_dataframe(\n                dataframe=df,\n                directory=valid_path,\n                x_col = 'id',\n                y_col = 'label',\n                subset='validation', \n                target_size=(96, 96),\n                batch_size=128,\n                shuffle=False,\n                class_mode='binary'\n                )","metadata":{"execution":{"iopub.status.busy":"2023-05-29T21:48:29.363699Z","iopub.execute_input":"2023-05-29T21:48:29.364412Z","iopub.status.idle":"2023-05-29T21:51:53.501704Z","shell.execute_reply.started":"2023-05-29T21:48:29.364375Z","shell.execute_reply":"2023-05-29T21:51:53.500659Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Building Model Architectures","metadata":{}},{"cell_type":"markdown","source":"### Classic CNN architectural design choices include:\n\n*  ***Convolutions use filters to transform or extract features from the image/pixels by sliding through the image and applying some matrix operations at each level. Since convolutions are an abstracting concept most of the design choices are surrounding the filters(kernels) themselves.***\n\n* ***Filter(kernel) size over the years has mainly converged to a 3x3 by the community. This seems to be very efficient as well as effective, however in some cases other sizes have been used in parallel with this window size.***\n\n* ***Initial layer feature extractions are more broad and deeper layers become increasingly granular. This is so the model can learn different degrees of patterns in the image. This is done by using the filter size and number of filters. Since after experimenting I have settled on using solely a 3x3 size filter, you will notice this being done by number of filters at each layer i.e. the `filters` parameter increasing throughout.***\n\n* ***Padding is used so as to preserve as much information as possible with each kernel. Without padding a 3x3 filter(kernel) would mean the outer edge of pixels would be lost. A 5x5 filter(kernel) would lose the outside 2 pixels around the entire image and so on.***\n\n* ***Striding may be used to reduce complexity and simplify the features the model uses. Striding is exactly what it sounds like, when the filter(kernel) traverses across and down the image it can do so one pixel at a time`(stride=1)`, or in larger strides skipping some pixels and abstracting away pixels in between along the way. This will significantly decrease the size/number of features/weights and make the model simpler and computation much less intensive for training/inference. The trade off would be abstracting away important features/weights and losing accuracy overall.***\n\n* ***Pooling can be used similar to striding wherein a specified number of pixels is abstracted into the max value in it's pool. This value pertains to the RGB value so that for example in the case of Max Pooling, all pixels in each particular section(pool size) of the image will be transformed to the same value in this case the max(in most cases darkest or strongest value if inverted) value for that region. The pool slides across and down the image applying this transformation just like the convolution filter/kernel.***\n\n* ***Dropout can be used to keep a random set of weights hidden(set to zero) so that the model must optimize using less weights and therefore will typically generalize better at inference time. We can apply dropout after any set of layers but it is typically done directly before or after a subsampling/abstracting layer such as a pooling layer. We can choose the percentage of weights to keep hidden.***\n\n* ***Regularization is when we transform the data to keep the values in the same range and typically limit the range to [0, 1] for example. When Dropout is being used regularization may not be unnecessary.***\n\n### Other Hyperparameters:\n\n* ***Activation functions are used to introduce non-linearity into the neural network so that it may better learn more complex pattersn. In this case I am using rectified linear units for most of the hidden layers with exponential linear units for some of the deeper layers. When we get to the output(classification) layer we will be using a sigmoid activation function which is well suited for seperation of binary classes.***\n\n* \n","metadata":{}},{"cell_type":"markdown","source":"### Model 1","metadata":{}},{"cell_type":"code","source":"with tf.device('/gpu:0'):  \n    model = Sequential()\n    \n    model.add(Conv2D(filters=16, kernel_size=3, padding='same', activation='relu', input_shape = (96, 96, 3)))\n    model.add(Conv2D(filters=16, kernel_size=3, padding='same', activation='relu'))\n    model.add(Conv2D(filters=16, kernel_size=3, padding='same', activation='relu'))\n    model.add(Dropout(0.3)) ## Hide some weights until inference to generalize better\n    model.add(MaxPooling2D(pool_size=3)) ## Subsample layer\n\n\n    model.add(Conv2D(filters=32, kernel_size=3, padding='same', activation='relu'))\n    model.add(Conv2D(filters=32, kernel_size=3, padding='same', activation='relu'))\n    model.add(Conv2D(filters=32, kernel_size=3, padding='same', activation='relu'))\n    model.add(Dropout(0.3)) ## Hide some weights until inference to generalize better\n    model.add(MaxPooling2D(pool_size=2)) ## Subsample layer\n\n\n    model.add(Conv2D(filters=64, kernel_size=3, padding='same', activation='relu'))\n    model.add(Conv2D(filters=64, kernel_size=3, padding='same', activation='relu'))\n    model.add(Conv2D(filters=64, kernel_size=3, padding='same', activation='relu'))\n    model.add(Dropout(0.3)) ## Hide some weights until inference to generalize better\n    model.add(MaxPooling2D(pool_size=2)) ## Subsample layer\n\n\n    model.add(Conv2D(filters=128, kernel_size=3, padding='same', activation='elu'))\n    model.add(Conv2D(filters=128, kernel_size=3, padding='same', activation='elu'))\n    model.add(Conv2D(filters=128, kernel_size=3, padding='same', activation='elu'))\n    model.add(Dropout(0.3)) ## Hide some weights until inference to generalize better\n    \n    ## This is the actual 'classifying' layer\n    model.add(Flatten()) ## Features into 1d array\n    model.add(Dense(512, activation='elu'))\n    model.add(Dropout(0.3)) ## Hide some weights until inference to generalize better\n    model.add(Dense(1, activation='sigmoid'))\n    model.summary()\n    ## Compile model and initialize hyperparameters for learning rate and evaluation metrics\n    model.compile(optimizer='adam', loss='binary_crossentropy', metrics=['accuracy', tf.keras.metrics.AUC()])\n\n    STEP_SIZE_TRAIN = train_generator.n//train_generator.batch_size\n    STEP_SIZE_VALID = validation_generator.n//validation_generator.batch_size\n\n    history = model.fit(train_generator,\n                        steps_per_epoch=STEP_SIZE_TRAIN,\n                        epochs=6,\n                        validation_data=validation_generator,\n                        validation_steps=STEP_SIZE_VALID)\n","metadata":{"execution":{"iopub.status.busy":"2023-05-29T23:06:26.631677Z","iopub.execute_input":"2023-05-29T23:06:26.632110Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"***Model 1 was my top performing custom CNN model and is a basic VGG style CNN. If I had more GPU resources I was going to apply the following to try for better results:***\n* ***Continue to increase the final dimension output before classification layer from `128`===>`256`===>`512`. With everything else the same I increased this from `32`===>`64`===>`128` and had an increase in performance each time. It is not a far fetched estimation that increasing further would continue this trend. This gives the classifier more features to train with. This is done in the Dense layer directly following the flatten layer.***\n\n* ***Try different methods for sub sampling layers. I settled on a Max Pooling method using a `3x3` sized pool window after trying different sized pools. I would like to try Average Pooling although not typically used for this data and problem would still be interesting to try. Another known technique for increased accuracy is using a kernel layers with striding to act as sub sampling. This abstracts away features just as pooling but preserves the values unlike pooling methods.***\n\n* ***Try removing the output layer of the neural network and taking features/weights from CNN layers and using as input features in a random forest or gradient boosted algorithm such as XGBoost for classification. See how results differ.***","metadata":{}},{"cell_type":"markdown","source":"### Model 2","metadata":{}},{"cell_type":"code","source":"# with tf.device('/gpu:0'):  \n#     model = Sequential()\n    \n#     model.add(Conv2D(filters=16, kernel_size=3, padding='same', activation='relu', input_shape = (96, 96, 3)))\n#     model.add(Conv2D(filters=16, kernel_size=3, padding='same', activation='relu'))\n#     model.add(Conv2D(filters=16, kernel_size=3, padding='same', activation='relu'))\n#     model.add(Dropout(0.3)) ## Regularization: hide some weights until inference to generalize better\n#     model.add(Conv2D(filters=32, kernel_size=3, strides=(3, 3), padding='same', activation='elu')) ## Try for more accurate results\n# #     model.add(Dropout(0.3)) ## Regularization: hide some weights until inference to generalize better\n# #     model.add(MaxPooling2D(pool_size=2)) ## Subsample layer\n\n\n#     model.add(Conv2D(filters=32, kernel_size=3, padding='same', activation='relu'))\n#     model.add(Conv2D(filters=32, kernel_size=3, padding='same', activation='relu'))\n#     model.add(Conv2D(filters=32, kernel_size=3, padding='same', activation='relu'))\n#     model.add(Dropout(0.3)) ## Regularization: hide some weights until inference to generalize better\n#     model.add(Conv2D(filters=64, kernel_size=3, strides=(3, 3), padding='same', activation='elu')) ## Try for more accurate results\n# #     model.add(Dropout(0.3)) ## Regularization: hide some weights until inference to generalize better\n# #     model.add(MaxPooling2D(pool_size=2)) ## Subsample layer\n\n\n#     model.add(Conv2D(filters=64, kernel_size=3, padding='same', activation='relu'))\n#     model.add(Conv2D(filters=64, kernel_size=3, padding='same', activation='relu'))\n#     model.add(Conv2D(filters=64, kernel_size=3, padding='same', activation='relu'))\n#     model.add(Dropout(0.3)) ## Regularization: hide some weights until inference to generalize better\n#     model.add(Conv2D(filters=128, kernel_size=3, strides=(3, 3), padding='same', activation='elu')) ## Try for more accurate results\n# #     model.add(Dropout(0.3)) ## Regularization: hide some weights until inference to generalize better\n# #     model.add(MaxPooling2D(pool_size=2)) ## Subsample layer\n\n\n#     model.add(Conv2D(filters=128, kernel_size=3, padding='same', activation='elu'))\n#     model.add(Conv2D(filters=128, kernel_size=3, padding='same', activation='elu'))\n#     model.add(Conv2D(filters=128, kernel_size=3, padding='same', activation='elu'))\n#     model.add(Dropout(0.3)) ## Regularization: hide some weights until inference to generalize better\n    \n#     ## This is the actual 'classifying' layer\n#     model.add(Flatten())\n#     model.add(Dense(128, activation='relu'))\n#     model.add(Dropout(0.3)) ## Regularization: hide some weights until inference to generalize better\n#     model.add(Dense(1, activation='sigmoid'))\n#     model.summary()\n#     ## Compile model and initialize hyperparameters for learning rate and evaluation metrics\n#     model.compile(optimizer='adam', loss='binary_crossentropy', metrics=['accuracy', tf.keras.metrics.AUC()])\n\n#     STEP_SIZE_TRAIN = train_generator.n//train_generator.batch_size\n#     STEP_SIZE_VALID = validation_generator.n//validation_generator.batch_size\n\n#     history = model.fit(train_generator,\n#                         steps_per_epoch=STEP_SIZE_TRAIN,\n#                         epochs=10,\n#                         validation_data=validation_generator,\n#                         validation_steps=STEP_SIZE_VALID)\n","metadata":{"execution":{"iopub.status.busy":"2023-05-29T22:15:49.005302Z","iopub.execute_input":"2023-05-29T22:15:49.005674Z","iopub.status.idle":"2023-05-29T22:15:49.012885Z","shell.execute_reply.started":"2023-05-29T22:15:49.005644Z","shell.execute_reply":"2023-05-29T22:15:49.012002Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"model.evaluate_generator(generator=validation_generator)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# summarize history for accuracy\nplt.plot(history.history['accuracy'])\nplt.plot(history.history['val_accuracy'])\nplt.title('model accuracy')\nplt.ylabel('accuracy')\nplt.xlabel('epoch')\nplt.legend(['train', 'validation'], loc='upper left')\nplt.show()\n\n# summarize history for loss\nplt.plot(history.history['loss'])\nplt.plot(history.history['val_loss'])\nplt.title('model loss')\nplt.ylabel('loss')\nplt.xlabel('epoch')\nplt.legend(['train', 'validation'], loc='upper left')\nplt.show()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Save the model\nmodel.save('Keras_CNN_4.h5')  ","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Inference","metadata":{}},{"cell_type":"code","source":"with tf.device('/GPU:0'):\n#     model = load_model('/kaggle/input/cancer-detections-cnn2/Keras_CNN_1.h5')\n    model = load_model('/kaggle/working/Keras_CNN_4.h5')\n\ntest_df = pd.read_csv(path + 'sample_submission.csv')\n\nTESTING_BATCH_SIZE = 64\ntesting_files = glob(os.path.join(test_path,'*.tif'))\nsubmission = pd.DataFrame()\nprint(len(testing_files))\n\nfor index in range(0, len(testing_files), TESTING_BATCH_SIZE):\n    data_frame = pd.DataFrame({'path': testing_files[index:index+TESTING_BATCH_SIZE]})\n    data_frame['id'] = data_frame.path.map(lambda x: x.split('/')[4].split(\".\")[0])\n    data_frame['image'] = data_frame['path'].map(imread)\n    \n    # Stack images into a batch\n    images = np.stack(data_frame.image, axis=0) / 255.0\n    # Predict on the entire batch\n    predictions = model.predict(images).flatten()\n    \n    data_frame['label'] = predictions\n    submission = pd.concat([submission, data_frame[[\"id\", \"label\"]]])","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Submission","metadata":{}},{"cell_type":"code","source":"submission.to_csv('submission.csv', index=False, header=True)\nprint(len(submission))\nsubmission.head()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}