{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.13","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"gpu","dataSources":[{"sourceId":11848,"databundleVersionId":862157,"sourceType":"competition"}],"dockerImageVersionId":30746,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":true}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Histopathologic Cancer Detection Project\n\nThe primary goal of this model was to develop a working predictive cnn that is capable of determining whether an image contains tumorous cells or is tumor free. The below results throughout the model are due to several iterations of trial and error to determine the ideal tuning parameters. While there is still a substantial amount of improvement to be made the end result is moving in the right direction of what a similar model can achieve.","metadata":{}},{"cell_type":"code","source":"##Prepare Environment\n\nimport numpy as np\nimport pandas as pd\nimport matplotlib.pyplot as plt\nimport matplotlib.image as mpimg\nimport pickle\nimport optree\n## commenting out to see if this is causing the keras error\n## import keras\n\n##### Moved before importing tensorflow to avoid errors\n##### Below block may need to be uncommented if necessary or errors are occuring for tensorflow\nimport os\n### os.environ['TS_CPP_MIN_LOG_LEVEL'] = '3'\n\nimport tensorflow as tf\n\nfrom sklearn.model_selection import train_test_split\nfrom tensorflow.keras.models import Sequential\nfrom tensorflow import keras\nfrom tensorflow.keras.layers import *\nfrom tensorflow.keras.preprocessing.image import ImageDataGenerator\nfrom tensorflow.keras import backend as k","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2024-07-28T20:08:18.429319Z","iopub.execute_input":"2024-07-28T20:08:18.430198Z","iopub.status.idle":"2024-07-28T20:08:18.436610Z","shell.execute_reply.started":"2024-07-28T20:08:18.430166Z","shell.execute_reply":"2024-07-28T20:08:18.435581Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Helping Functions\n\nThese functions were created to help in creating a combined history with the results of each iteration of epochs in the below model. Thanks to this method the results of subsequent iterations can be merged together to generate a running history when creating the plot to show the models results.","metadata":{}},{"cell_type":"code","source":"## Helper Functions\n\ndef merge_history(hlist):\n    history = {}\n    for k in hlist[0].history.keys():\n        history[k] = sum([h.history[k] for h in hlist],[])\n    return history\n\ndef vis_training(h, start=1):\n    epoch_range = range(start, len(h['loss'])+1)\n    s = slice(start-1, None)\n    \n    plt.figure(figsize=[14,4])\n    \n    n = int(len(h.keys()) / 2)\n    \n    for i in range(n):\n        k = list(h.keys())[i]\n        plt.subplot(1,n,i+1)\n        plt.plot(epoch_range, h[k][s], label = 'Training')\n        plt.plot(epoch_range, h['val_' + k][s], label = 'Validation')\n        plt.xlabel('Epoch'); plt.ylabel(k); plt.title(k)\n        plt.grid()\n        plt.legend()\n        \n    plt.tight_layout()\n    plt.show()","metadata":{"execution":{"iopub.status.busy":"2024-07-28T19:52:16.413191Z","iopub.execute_input":"2024-07-28T19:52:16.413905Z","iopub.status.idle":"2024-07-28T19:52:16.422992Z","shell.execute_reply.started":"2024-07-28T19:52:16.413870Z","shell.execute_reply":"2024-07-28T19:52:16.421961Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Importing and Evaluating Data\n\nThe below setion is created to import and evaluate the data being used in this project to insure that the information is accurate to what is expecting and to determine what route would be best to continue forward in model selection and adaptation.","metadata":{}},{"cell_type":"code","source":"## Importing training data for testing - Set to string since Categorical Classes are being used. \n## Avoiding Binary in case of additional classes being added in the future.\ntrain = pd.read_csv('../input/histopathologic-cancer-detection/train_labels.csv', dtype=str)\nprint(train.shape)","metadata":{"execution":{"iopub.status.busy":"2024-07-28T19:52:16.424265Z","iopub.execute_input":"2024-07-28T19:52:16.424628Z","iopub.status.idle":"2024-07-28T19:52:16.794843Z","shell.execute_reply.started":"2024-07-28T19:52:16.424595Z","shell.execute_reply":"2024-07-28T19:52:16.793915Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"## Concatenating train.id with .tif to ensure model can properly read the images.\ntrain.id = train.id + '.tif'\ntrain.head(10)","metadata":{"execution":{"iopub.status.busy":"2024-07-28T19:52:16.797039Z","iopub.execute_input":"2024-07-28T19:52:16.797330Z","iopub.status.idle":"2024-07-28T19:52:16.844291Z","shell.execute_reply.started":"2024-07-28T19:52:16.797305Z","shell.execute_reply":"2024-07-28T19:52:16.843402Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"## Label Distribution\n\n(train.label.value_counts() / len(train)).to_frame().sort_index().T","metadata":{"execution":{"iopub.status.busy":"2024-07-28T19:52:16.845448Z","iopub.execute_input":"2024-07-28T19:52:16.845787Z","iopub.status.idle":"2024-07-28T19:52:16.889159Z","shell.execute_reply.started":"2024-07-28T19:52:16.845755Z","shell.execute_reply":"2024-07-28T19:52:16.888330Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Image Extraction and Review\n\nBy extracting the images from the data imported it can be confirmed if the images are loading correctly through visual confirmation. Once the images are successfully loaded they can be further analyzed to determine if different techniques may be necessary to improve the model in the future. Things like image augmentation for instance may become much more relevant if the sample images show a lot of variation in format or placement of the cells within the pixel boundaries.","metadata":{}},{"cell_type":"code","source":"## Image Extraction\n\ntrain_path = \"../input/histopathologic-cancer-detection/train\"\n\nsample = train.sample(n=16).reset_index()\n\nplt.figure(figsize=(6,6))\n\nfor i, row in sample.iterrows():\n    \n    img = mpimg.imread(f'../input/histopathologic-cancer-detection/train/{row.id}')\n    label = row.label\n    \n    plt.subplot(4,4,i+1)\n    plt.imshow(img)\n    plt.text(0,-5, f'Class {label}', color='k')\n    plt.axis('off')\n    \nplt.tight_layout()\nplt.show()\n    ","metadata":{"execution":{"iopub.status.busy":"2024-07-28T20:51:37.091988Z","iopub.execute_input":"2024-07-28T20:51:37.092357Z","iopub.status.idle":"2024-07-28T20:51:38.274310Z","shell.execute_reply.started":"2024-07-28T20:51:37.092325Z","shell.execute_reply":"2024-07-28T20:51:38.273473Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Data Splitting - Creating Training & Testing Sets\n\nSplitting of the data below into training and testing data sets was conducted through random selection and split 80/20 to insure testing can be conducted on a smaller selection while maintaining some data without changes. By creating two separate data frames for testing and training the original data will remain unaltered to maintain a level of control. \n\nWithin this section the image data generators are loaded and any image augmentation can be applied at this point. Due to time constraints on a previous version of this model the image augmentation was commented out to focus on reducing unstable results first and can be added back in once stability is acceptable.","metadata":{}},{"cell_type":"code","source":"## Training and Validation Set Creation\ntrain_df, valid_df = train_test_split(train, test_size = 0.2, random_state = 1, stratify = train.label)\n\n## Data Generators\n## Image Augmentation set only for training to maintain controls.\n\ntrain_datagen = ImageDataGenerator(\n    rescale = 1/255,\n##    shear_range = 0.1,\n##    zoom_range = 0.1\n                                  )\nvalidation_datagen = ImageDataGenerator(\n    rescale = 1/255)","metadata":{"execution":{"iopub.status.busy":"2024-07-28T19:52:17.635705Z","iopub.execute_input":"2024-07-28T19:52:17.636037Z","iopub.status.idle":"2024-07-28T19:52:17.962808Z","shell.execute_reply.started":"2024-07-28T19:52:17.636008Z","shell.execute_reply":"2024-07-28T19:52:17.962045Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Creating Data Loaders\n\nEstablishing the parameters for how the images are to be loaded will allow the model to appropriately read the images. The original intent with BATCH_SIZE was to create a variable that could be adjusted in all other instances when adjustments were to be made. \n\nUnfortunately an error was occuring when trying to feed the variable in so a hard coded value was used to solve this issue. In doing so it is important in testing to adjust all batch_size instances throughout the model to insure accuracy when running.\n\nNote: By selecting the class_mode = 'categorical' there is some flexibility in the model to introduce alternative results other than a \"Yes\" or \"No\" given the validation set is limited to a 0 or 1 but if another validation set was established with say a third or more option stating \"questionable\" instead of just being a tumor is or is not present then the model would be capable of having additional results. This prevents the model from being stuck to only predicting binary results.","metadata":{}},{"cell_type":"code","source":"## Setting Batch Sizes - Attempting manual batch_size to fix an error\n## BATCH_SIZE = 32 (changed from 64) \n\n## Setting Batch Size with variable, seed, and target size.\ntrain_loader = train_datagen.flow_from_dataframe(\n    dataframe = valid_df,\n    directory = train_path,\n    x_col = 'id',\n    y_col = 'label',\n    batch_size = 96,\n    seed = 1,\n    shuffle = True,\n    class_mode = 'categorical',\n    target_size = (96,96)\n)\n\n\nvalid_loader = train_datagen.flow_from_dataframe(\n    dataframe = valid_df,\n    directory = train_path,\n    x_col = 'id',\n    y_col = 'label',\n    batch_size = 96,\n    seed = 1,\n    shuffle = True,\n    class_mode = 'categorical',\n    target_size = (96,96)\n)","metadata":{"execution":{"iopub.status.busy":"2024-07-28T21:19:35.944782Z","iopub.execute_input":"2024-07-28T21:19:35.945177Z","iopub.status.idle":"2024-07-28T21:19:57.909472Z","shell.execute_reply.started":"2024-07-28T21:19:35.945148Z","shell.execute_reply":"2024-07-28T21:19:57.908499Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Assessing Training and Validation Steps\n\nInitially I had not fully understood this portion of coding when I had seen it implemented in other's submissions. However, it become apparent that there is quite a bit of value in knowing the step counts as this gives a more clear idea of what batch sizes or epoch sizes will be optimal when making adjustments to the model.","metadata":{}},{"cell_type":"code","source":"## Verify number of steps is correct according to previous parameters\nTR_STEPS = len(train_loader)\nVA_STEPS = len(valid_loader)\n\nprint(TR_STEPS) \nprint(VA_STEPS)","metadata":{"execution":{"iopub.status.busy":"2024-07-28T21:21:06.587326Z","iopub.execute_input":"2024-07-28T21:21:06.588023Z","iopub.status.idle":"2024-07-28T21:21:06.592994Z","shell.execute_reply.started":"2024-07-28T21:21:06.587991Z","shell.execute_reply":"2024-07-28T21:21:06.591984Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Introducing Pre-Training Base Model\n\nThrough the introduction of a pre-trained model the overall performance of the model improved greatly. Having a baseline model to adapt allows for much easier tweaking rather than starting from scratch and being unsure of performance with each adjustment. By assigning similar parameters in the base model to the parameters in the image data generators this will reduce errors and allow for the pre-trained model to be used as a base model to build off of.","metadata":{}},{"cell_type":"code","source":"## Pre-Trained Model to set Convolutional Base\n## Setting image size to 32x32 to reduce initial load time.\n## Creating Base Model\n\nbase_model = tf.keras.applications.VGG19(\n    input_shape = (96,96,3),\n    include_top = False,\n    weights = 'imagenet'\n)\n\nbase_model.trainable = False\n\n## Verify Base Model looks correct and layers show appropriately\nbase_model.summary()","metadata":{"execution":{"iopub.status.busy":"2024-07-28T21:21:09.598240Z","iopub.execute_input":"2024-07-28T21:21:09.598600Z","iopub.status.idle":"2024-07-28T21:21:10.072813Z","shell.execute_reply.started":"2024-07-28T21:21:09.598572Z","shell.execute_reply":"2024-07-28T21:21:10.071989Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Constructing Convultional Nueral Network\n\nThrough the implementation of batch normalization and flattening the model will be able to use a more condensed scale to generate results and will prevent extreme outliers that would cause the accuracy of the model to decrease tremendously. Having lower and upper bounds limited will insure that those extremes cannot create false averages or \"volatile\" predictions that will throw off other accurate predictions.","metadata":{}},{"cell_type":"code","source":"## Building Convolutional Neural Network\nnp.random.seed(1)\ntf.random.set_seed(1)\n\ncnn = Sequential([\n    base_model,\n    BatchNormalization(),\n    \n    Flatten(),\n    \n    Dense(96, activation = 'relu'),\n    Dropout(0.5),\n    Dense(64, activation = 'relu'),\n    Dropout(0.5),\n    BatchNormalization(),\n    Dense(2, activation = 'softmax')\n])\n","metadata":{"execution":{"iopub.status.busy":"2024-07-28T21:21:45.015052Z","iopub.execute_input":"2024-07-28T21:21:45.015932Z","iopub.status.idle":"2024-07-28T21:21:45.049493Z","shell.execute_reply.started":"2024-07-28T21:21:45.015896Z","shell.execute_reply":"2024-07-28T21:21:45.048562Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"## Causing following error: ValueError: Undefined shapes are not supported.\n## cnn.summary()","metadata":{"execution":{"iopub.status.busy":"2024-07-28T21:13:26.729278Z","iopub.execute_input":"2024-07-28T21:13:26.729855Z","iopub.status.idle":"2024-07-28T21:13:26.733604Z","shell.execute_reply.started":"2024-07-28T21:13:26.729813Z","shell.execute_reply":"2024-07-28T21:13:26.732716Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Initial Optimizer Settings","metadata":{}},{"cell_type":"code","source":"## Fine tuning - setting epochs lower to allow for quicker testing at first.\n## WEEK 6 - Added additional 0 to optimizer to see if results improve with large images.\n\nopt = tf.keras.optimizers.Adam(0.0001)\ncnn.compile(loss='categorical_crossentropy', optimizer = opt, metrics=['accuracy', tf.keras.metrics.AUC()])","metadata":{"execution":{"iopub.status.busy":"2024-07-28T21:21:49.657496Z","iopub.execute_input":"2024-07-28T21:21:49.658161Z","iopub.status.idle":"2024-07-28T21:21:49.688268Z","shell.execute_reply.started":"2024-07-28T21:21:49.658128Z","shell.execute_reply":"2024-07-28T21:21:49.687324Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Establishing History and First Iteration of Epochs\n\nHere the functions for establishing and merging history are put into use. As the model is fit the end results can then be set within a collective history so the progress of the models learning can be monitored. Each iteration of epochs can then be merged and a smooth plot of each new set of epochs can be seen together.\n\nNote: steps_per_epoch and validation_steps were also generating errors and were commented out. This appeared to be due to conflicts as the function was already pulling this information and having the parameter set again only caused conflicts.","metadata":{}},{"cell_type":"code","source":"## First iteration to test 16 Epochs using parameters set.\n## Ommitting steps_per_epoch and validation_steps correcting error in epoch processing. <--- Error Fix.\n\nh1 = cnn.fit(\n    x = train_loader,\n##    steps_per_epoch = TR_STEPS,\n    validation_data = valid_loader,\n##    validation_steps = VA_STEPS,\n    epochs = 30,\n    verbose = 1\n)","metadata":{"execution":{"iopub.status.busy":"2024-07-28T21:29:32.556261Z","iopub.execute_input":"2024-07-28T21:29:32.556940Z","iopub.status.idle":"2024-07-28T22:02:16.977918Z","shell.execute_reply.started":"2024-07-28T21:29:32.556906Z","shell.execute_reply":"2024-07-28T22:02:16.977128Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"## Merge history to store results\nhistory = merge_history([h1])\nvis_training(history)","metadata":{"execution":{"iopub.status.busy":"2024-07-28T22:04:16.390748Z","iopub.execute_input":"2024-07-28T22:04:16.391643Z","iopub.status.idle":"2024-07-28T22:04:17.265752Z","shell.execute_reply.started":"2024-07-28T22:04:16.391609Z","shell.execute_reply":"2024-07-28T22:04:17.264888Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Second Iteration with Learning Rate Schedule.\n\nWith the learning rate schedule applied the model was able to adjust how the optimizer was being applied through the steps to try to allow for more finite adjustments and to prevent unstable changes in both directions. By exponentially decaying the learning rate the predictions will becoming less volatile as the steps continue to try to eliminate large swings in the results. Over time this generates smoother predictions rather than generating large jumps.","metadata":{}},{"cell_type":"code","source":"## Fine tuning for 2nd iteration - Check learning rate schedules for potential model improvements.\nbase_model.trainable = False\nlearning_rate = tf.Variable(0.0001, trainable = False)\n\nlearning_rate_schedule = tf.keras.optimizers.schedules.ExponentialDecay(\n    initial_learning_rate = 0.001,\n    decay_steps = 10000,\n    decay_rate = 0.96\n)\n\nopt2 = tf.keras.optimizers.Adam(learning_rate=learning_rate_schedule)\n## k.set_value(cnn.optimizer.learning_rate, 0.0001)\n\ncnn.compile(\n    loss = 'categorical_crossentropy',\n    optimizer = opt2,\n    metrics = ['accuracy', tf.keras.metrics.AUC()]\n)\n\n","metadata":{"execution":{"iopub.status.busy":"2024-07-28T22:21:30.262080Z","iopub.execute_input":"2024-07-28T22:21:30.262805Z","iopub.status.idle":"2024-07-28T22:21:30.280241Z","shell.execute_reply.started":"2024-07-28T22:21:30.262772Z","shell.execute_reply":"2024-07-28T22:21:30.279458Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"## Fine tuned using altered optimizer and learning rate scheduler to compare results of previous iteration to new h2 model.\n\nh2 = cnn.fit(\n    x = train_loader,\n    validation_data = valid_loader,\n    epochs = 30,\n    verbose = 1\n)","metadata":{"execution":{"iopub.status.busy":"2024-07-28T22:22:19.566051Z","iopub.execute_input":"2024-07-28T22:22:19.566383Z","iopub.status.idle":"2024-07-28T22:54:06.358551Z","shell.execute_reply.started":"2024-07-28T22:22:19.566358Z","shell.execute_reply":"2024-07-28T22:54:06.357607Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"## Merge h2 history with h1 history can display results.\n\nh2.history['auc'] = h2.history['auc_1']\nh2.history['val_auc'] = h2.history['val_auc_1']","metadata":{"execution":{"iopub.status.busy":"2024-07-28T23:05:58.530364Z","iopub.execute_input":"2024-07-28T23:05:58.531205Z","iopub.status.idle":"2024-07-28T23:05:58.535536Z","shell.execute_reply.started":"2024-07-28T23:05:58.531173Z","shell.execute_reply":"2024-07-28T23:05:58.534593Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"history = merge_history([h1, h2])\nvis_training(history, start = 30)","metadata":{"execution":{"iopub.status.busy":"2024-07-28T23:06:00.775197Z","iopub.execute_input":"2024-07-28T23:06:00.775872Z","iopub.status.idle":"2024-07-28T23:06:01.620477Z","shell.execute_reply.started":"2024-07-28T23:06:00.775820Z","shell.execute_reply":"2024-07-28T23:06:01.619564Z"}}},{"cell_type":"code","source":"## Third Iteration\n\nh3 = cnn.fit(\n    x = train_loader,\n    validation_data = valid_loader,\n    epochs = 30,\n    verbose = 1\n)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"## Merge h3 history with h2 and h1 history can display results.\n\nh3.history['auc'] = h3.history['auc_1']\nh3.history['val_auc'] = h3.history['val_auc_1']","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"history = merge_history([h1, h2, h3])\nvis_training(history, start = 30)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"h4 = cnn.fit(\n    x = train_loader,\n    validation_data = valid_loader,\n    epochs = 30,\n    verbose = 1\n)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"## Assign and merge 4th iteration to verify results after error.\n\nh4.history['auc'] = h4.history['auc_1']\nh4.history['val_auc'] = h4.history['val_auc_1']","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Final Iteration and Collective Results\n\nWith the fourth iteration completed the final vis_training will display the results of the four sets of epochs. This final graph will give a good indication as to whether the training set was appropriately learning and creating predictions in comparison to the validation set.","metadata":{}},{"cell_type":"code","source":"history = merge_history([h1, h2, h3, h4])\nvis_training(history, start = 30)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Saving and Generating Submission\n\nNow that the model has been checked the below section saves the model and is applied to the test data set. The end results can then be submitted to determine if the model is a good fit for the competition.","metadata":{}},{"cell_type":"code","source":"## Saving output\n\ncnn.save('HCD.h5')\npickle.dump(history, open(f'HCD.pk1', 'wb'))","metadata":{"execution":{"iopub.status.busy":"2024-07-28T23:11:05.278505Z","iopub.execute_input":"2024-07-28T23:11:05.279178Z","iopub.status.idle":"2024-07-28T23:11:05.496388Z","shell.execute_reply.started":"2024-07-28T23:11:05.279145Z","shell.execute_reply":"2024-07-28T23:11:05.495611Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"## Setting up for submission\ntest = pd.read_csv('../input/histopathologic-cancer-detection/sample_submission.csv')\ntest_path = \"../input/histopathologic-cancer-detection/test\"\ntest['filename'] = test.id + '.tif'","metadata":{"execution":{"iopub.status.busy":"2024-07-28T23:20:00.804419Z","iopub.execute_input":"2024-07-28T23:20:00.805134Z","iopub.status.idle":"2024-07-28T23:20:00.883093Z","shell.execute_reply.started":"2024-07-28T23:20:00.805100Z","shell.execute_reply":"2024-07-28T23:20:00.882329Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_datagen = ImageDataGenerator(rescale=1/255)\n\ntest_loader = train_datagen.flow_from_dataframe(\n    dataframe = test,\n    directory = test_path,\n    x_col = 'filename',\n    batch_size = 96,\n    shuffle = False,\n    class_mode = None,\n    target_size = (96,96)\n)","metadata":{"execution":{"iopub.status.busy":"2024-07-28T23:20:02.557784Z","iopub.execute_input":"2024-07-28T23:20:02.558436Z","iopub.status.idle":"2024-07-28T23:22:00.332929Z","shell.execute_reply.started":"2024-07-28T23:20:02.558407Z","shell.execute_reply":"2024-07-28T23:22:00.332041Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_probs = cnn.predict(test_loader)\ntest_pred = np.argmax(test_probs, axis = 1)\n","metadata":{"execution":{"iopub.status.busy":"2024-07-28T23:22:11.612016Z","iopub.execute_input":"2024-07-28T23:22:11.612486Z","iopub.status.idle":"2024-07-28T23:26:55.443221Z","shell.execute_reply.started":"2024-07-28T23:22:11.612452Z","shell.execute_reply":"2024-07-28T23:26:55.441565Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"submission = pd.read_csv('../input/histopathologic-cancer-detection/sample_submission.csv')\nsubmission.head()","metadata":{"execution":{"iopub.status.busy":"2024-07-28T23:27:24.037136Z","iopub.execute_input":"2024-07-28T23:27:24.037895Z","iopub.status.idle":"2024-07-28T23:27:24.093960Z","shell.execute_reply.started":"2024-07-28T23:27:24.037862Z","shell.execute_reply":"2024-07-28T23:27:24.093062Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"submission.label = test_probs[:,1]\nsubmission.head()","metadata":{"execution":{"iopub.status.busy":"2024-07-28T23:27:28.829971Z","iopub.execute_input":"2024-07-28T23:27:28.830838Z","iopub.status.idle":"2024-07-28T23:27:28.840045Z","shell.execute_reply.started":"2024-07-28T23:27:28.830792Z","shell.execute_reply":"2024-07-28T23:27:28.839162Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"submission.to_csv('submission.csv', header = True, index = False)","metadata":{},"execution_count":null,"outputs":[]}]}