{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.13","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"nvidiaTeslaT4","dataSources":[{"sourceId":11848,"databundleVersionId":862157,"sourceType":"competition"}],"dockerImageVersionId":30684,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":true}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"## Abstract\n\nEarly detection of metastatic cancer in histopathologic scans is pivotal for effective treatment planning and improved patient outcomes. This project aims to advance the field of digital pathology through the development of a robust deep learning algorithm for the binary classification of histopathologic image patches. Leveraging a modified and duplicate-free version of the PatchCamelyon (PCam) dataset, the project introduces a convolutional neural network (CNN)-based approach to discern metastatic from non-metastatic tissue samples. The MetaDetect AI system undergoes rigorous training, validation, and testing phases, employing data augmentation and normalization techniques to enhance the dataset's diversity and representativeness. With a focus on both the accuracy and interpretability of the model, the project incorporates evaluation metrics such as precision, recall, and the area under the ROC curve, alongside visualization techniques like saliency maps to ensure diagnostic reliability and trustworthiness. The project encompasses the entire spectrum of development stages, from initial data preprocessing to the intricacies of model training and comprehensive evaluation.","metadata":{}},{"cell_type":"code","source":"import numpy as np\nimport pandas as pd \nimport matplotlib.pyplot as plt\nimport sklearn\nimport keras\nimport tensorflow as tf\nfrom tensorflow.keras.preprocessing.image import ImageDataGenerator\nimport itertools\nimport shutil\nimport os\nimport gc\nimport cv2 \nfrom PIL import Image\nfrom sklearn.utils import shuffle\nfrom sklearn.metrics import confusion_matrix\nfrom sklearn.model_selection import train_test_split\nfrom tensorflow.keras.layers import Conv2D, MaxPooling2D\nfrom tensorflow.keras.layers import Dense, Dropout, Flatten, Activation\nfrom tensorflow.keras.models import Sequential\nfrom tensorflow.keras.callbacks import EarlyStopping, ReduceLROnPlateau, ModelCheckpoint\nfrom tensorflow.keras.optimizers import Adam","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2024-04-14T18:34:18.051711Z","iopub.execute_input":"2024-04-14T18:34:18.052294Z","iopub.status.idle":"2024-04-14T18:34:33.814656Z","shell.execute_reply.started":"2024-04-14T18:34:18.052265Z","shell.execute_reply":"2024-04-14T18:34:33.813812Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Exploratory Data Analysis (EDA)\n\n### Image Size and Features\n- The dataset is composed of histopathological images, each with a resolution of 96x96 pixels.\n- The label of interest is determined by the central region of the image (32x32 pixels). A positive label is assigned if this central region contains at least one pixel of tumor tissue, which is indicative of metastatic cancer.\n- The dataset is structured to support fully-convolutional models without the need for zero-padding. This ensures that the model's behavior remains consistent when scaled up to analyze whole-slide images.\n\n### Dataset Integrity\n- The original PCam dataset included duplicates due to probabilistic sampling methods. However, the version available on Kaggle has been curated to remove these duplicates, preserving the integrity of the dataset.\n- Despite the removal of duplicates, the Kaggle version retains the same data splits as the original PCam benchmark, ensuring consistency for model development and validation.\n\n### Label Distribution\n- Initially, it was believed that the training data would contain a 50/50 split between positive and negative labels. However, further analysis revealed a distribution closer to 60/40.\n- This imbalance in label distribution should be considered in the model design process to prevent bias in the prediction outcomes.\n\n### Data Relevance\n- The dataset merges two independent datasets from Radboud University Medical Center (Nijmegen, the Netherlands) and the University Medical Center Utrecht (Utrecht, the Netherlands).\n- These datasets are derived from routine clinical procedures and have been vetted by trained pathologists, affirming their applicability for training models aimed at identifying metastases.\n\n### Next Steps in Model Development\n- With the data exploration complete, the focus will now shift to designing and training the model that will accurately classify the histopathological images.\n","metadata":{}},{"cell_type":"code","source":"# Print list of files and directories in folder\ninput_dir = '/kaggle/input/histopathologic-cancer-detection'\nlist_l = [os.path.join(input_dir, x) for x in os.listdir(input_dir)]\nlist_l","metadata":{"execution":{"iopub.status.busy":"2024-04-14T18:34:33.816593Z","iopub.execute_input":"2024-04-14T18:34:33.817119Z","iopub.status.idle":"2024-04-14T18:34:33.825165Z","shell.execute_reply.started":"2024-04-14T18:34:33.817087Z","shell.execute_reply":"2024-04-14T18:34:33.824297Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_labels = pd.read_csv(list_l[1])\ndf_samples = pd.read_csv(list_l[0])\nprint(df_labels.shape)","metadata":{"execution":{"iopub.status.busy":"2024-04-14T18:34:33.826833Z","iopub.execute_input":"2024-04-14T18:34:33.827352Z","iopub.status.idle":"2024-04-14T18:34:34.324909Z","shell.execute_reply.started":"2024-04-14T18:34:33.827321Z","shell.execute_reply":"2024-04-14T18:34:34.323989Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(df_labels['label'].value_counts())\ntrain = list_l[3]\ntest = list_l[2]\n\nprint(\"Number of training images: {}\".format(len(os.listdir(train))))\nprint(\"Number of test images: {}\".format(len(os.listdir(test))))\n\nmissing_values = df_labels.isnull().sum()\nprint(\"Missing values: {}\".format( missing_values))","metadata":{"execution":{"iopub.status.busy":"2024-04-14T18:34:34.326869Z","iopub.execute_input":"2024-04-14T18:34:34.327156Z","iopub.status.idle":"2024-04-14T18:34:38.304636Z","shell.execute_reply.started":"2024-04-14T18:34:34.327124Z","shell.execute_reply":"2024-04-14T18:34:38.303637Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Get a list of training and test image filenames\nimg_train_filenames = os.listdir(train)\nimg_test_filenames = os.listdir(test)\nfig, axes = plt.subplots(2, 5, figsize=(10, 4))\n\nfor i in range(10): \n    row = i // 5\n    col = i % 5\n    ax = axes[row, col]\n    img_path = os.path.join(train, img_train_filenames[i])\n    img = Image.open(img_path)\n    img = img.convert('RGB') \n    label = df_labels.loc[df_labels['id'] == img_train_filenames[i].split('.')[0], 'label'].values[0]\n    ax.imshow(img)  \n    ax.set_title(f\"{i+1} - Label: {label}\")\n    ax.axis('off')  # Hide the axis\n\nplt.show()\n","metadata":{"execution":{"iopub.status.busy":"2024-04-14T18:34:38.305951Z","iopub.execute_input":"2024-04-14T18:34:38.306280Z","iopub.status.idle":"2024-04-14T18:34:39.836830Z","shell.execute_reply.started":"2024-04-14T18:34:38.306254Z","shell.execute_reply":"2024-04-14T18:34:39.835909Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Data Modeling Strategy\n\n## Sample Size Consideration\n\n- **Efficient Sampling**: Working with a smaller, randomized sample of the 220,000 available training images can significantly reduce processing time without substantially compromising the learning process. Defined sample # 70000, randomly selected\n\n## Model Design\n\nIn this section, we outline the architecture and configuration of our Convolutional Neural Network (CNN) model designed for binary classification of histopathological images. The model is implemented using TensorFlow and Keras.\n\n### Architecture\n\nThe CNN architecture consists of the following layers:\n- **Convolutional Layers**: Three sets of convolutional layers with ReLU activation functions to capture patterns and features from the images. Each set contains three convolutional operations followed by a ReLU activation.\n  - First layer filters: 32\n  - Second layer filters: 64\n  - Third layer filters: 128\n- **Max Pooling Layers**: Following each set of convolutional layers, a max pooling layer reduces the dimensionality of the data, which helps in reducing the computational cost and overfitting.\n- **Dropout Layers**: Dropout layers are included after each max pooling operation with a dropout rate of 0.3 to prevent overfitting by randomly dropping units from the neural network during training.\n- **Flatten Layer**: Converts the 2D feature maps to a 1D feature vector before passing them to the dense layers.\n- **Dense Layers**: A dense layer with 256 units and ReLU activation followed by a dropout layer with a rate of 0.3. The output layer uses a softmax activation function to output the probability distribution over the two classes.\n\n### Compilation\n\nThe model is compiled using the Adam optimizer with a learning rate of 0.0001. The loss function used is binary cross-entropy, which is appropriate for binary classification problems. Accuracy is the metric used to monitor the training and validation performance of the model.\n\n### Training\n\nThe model is trained using data augmented images to improve model robustness and to help generalize better when making predictions on new, unseen data. The training process includes callbacks such as `EarlyStopping`, `ReduceLROnPlateau`, and `ModelCheckpoint` to save the best model based on validation accuracy and to adjust the learning rate dynamically based on validation loss performance.\n\nThis design aims to efficiently learn discriminative features from histopathological images and accurately classify them into their respective categories, thus aiding in the diagnosis process.\n\n\n\n\n","metadata":{}},{"cell_type":"code","source":"IMAGE_SIZE=96\nIMAGE_CHANNELS=3\nSAMPLE_SIZE=70000  \n\ntr_0=df_labels[df_labels['label']==0].sample(SAMPLE_SIZE,random_state=101)\ntr_1=df_labels[df_labels['label']==1].sample(SAMPLE_SIZE,random_state=101)\n\n# concat the dataframes\ndf_sam = pd.concat([tr_0, tr_1], axis=0).reset_index(drop=True)\n# shuffle\ndf_sam = shuffle(df_sam)\n\ndf_sam['label'].value_counts()","metadata":{"execution":{"iopub.status.busy":"2024-04-14T18:34:39.837939Z","iopub.execute_input":"2024-04-14T18:34:39.838213Z","iopub.status.idle":"2024-04-14T18:34:39.903007Z","shell.execute_reply.started":"2024-04-14T18:34:39.838188Z","shell.execute_reply":"2024-04-14T18:34:39.902129Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"\ny = df_sam['label']\ndf_sam_train, df_sam_val = train_test_split(df_sam, test_size=0.10, random_state=101, stratify=y)\nprint(df_sam_train.shape)\nprint(df_sam_val.shape)\n# Get a list of train and val images\n\ntrain_list = list(df_sam_train['id'])\nval_list = list(df_sam_val['id'])","metadata":{"execution":{"iopub.status.busy":"2024-04-14T18:35:49.276263Z","iopub.execute_input":"2024-04-14T18:35:49.277025Z","iopub.status.idle":"2024-04-14T18:35:49.374178Z","shell.execute_reply.started":"2024-04-14T18:35:49.276993Z","shell.execute_reply":"2024-04-14T18:35:49.373260Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_sam.head()","metadata":{"execution":{"iopub.status.busy":"2024-04-14T18:35:52.048492Z","iopub.execute_input":"2024-04-14T18:35:52.049186Z","iopub.status.idle":"2024-04-14T18:35:52.062391Z","shell.execute_reply.started":"2024-04-14T18:35:52.049157Z","shell.execute_reply":"2024-04-14T18:35:52.061403Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Create a new directory so that we will be using the ImageDataGenerator\nbase_dir='base_dir'\nos.mkdir(base_dir)\n\n# create 2 folders inside 'base_dir':\n# dir_tr\n    # a_no_tumor_tissue\n    # b_has_tumor_tissue\n# dir_val\n    # a_no_tumor_tissue\n    # b_has_tumor_tissue\n# create a path to 'base_dir' to which we will join the names of the new folders\n# train_dir\ndir_tr = os.path.join(base_dir, 'dir_tr')\nos.mkdir(dir_tr)\n# val_dir\ndir_val = os.path.join(base_dir, 'dir_val')\nos.mkdir(dir_val)\n\n# [CREATE FOLDERS INSIDE THE TRAIN AND VALIDATION FOLDERS]\n\n# create new folders inside train_dir\nno_tumor_tissue = os.path.join(dir_tr, 'a_no_tumor_tissue')\nos.mkdir(no_tumor_tissue)\nhas_tumor_tissue = os.path.join(dir_tr, 'b_has_tumor_tissue')\nos.mkdir(has_tumor_tissue)\n\n# create new folders inside val_dir\nno_tumor_tissue = os.path.join(dir_val, 'a_no_tumor_tissue')\nos.mkdir(no_tumor_tissue)\nhas_tumor_tissue = os.path.join(dir_val, 'b_has_tumor_tissue')\nos.mkdir(has_tumor_tissue)","metadata":{"execution":{"iopub.status.busy":"2024-04-14T18:35:54.254988Z","iopub.execute_input":"2024-04-14T18:35:54.255795Z","iopub.status.idle":"2024-04-14T18:35:54.268121Z","shell.execute_reply.started":"2024-04-14T18:35:54.255767Z","shell.execute_reply":"2024-04-14T18:35:54.267123Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"\ndf_sam.set_index('id', inplace=True)\n\n# Transfer the train images\n\nfor image in train_list:\n    \n    # the id in the csv file does not have the .tif extension therefore we add it here\n    fname = image + '.tif'\n    # get the label for a certain image\n    target = df_sam.loc[image,'label']\n    \n    # these must match the folder names\n    if target == 0:\n        label = 'a_no_tumor_tissue'\n    if target == 1:\n        label = 'b_has_tumor_tissue'\n    \n    # source path to image\n    src = os.path.join(list_l[3], fname)\n    # destination path to image\n    dst = os.path.join(dir_tr, label, fname)\n    # copy the image from the source to the destination\n    shutil.copyfile(src, dst)\n\n\n","metadata":{"execution":{"iopub.status.busy":"2024-04-14T18:35:58.366600Z","iopub.execute_input":"2024-04-14T18:35:58.367285Z","iopub.status.idle":"2024-04-14T18:53:27.924972Z","shell.execute_reply.started":"2024-04-14T18:35:58.367251Z","shell.execute_reply":"2024-04-14T18:53:27.923809Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# check how many train images we have in each folder\n\nprint(len(os.listdir('base_dir/dir_tr/a_no_tumor_tissue')))\nprint(len(os.listdir('base_dir/dir_tr/b_has_tumor_tissue')))","metadata":{"execution":{"iopub.status.busy":"2024-04-14T18:53:27.926773Z","iopub.execute_input":"2024-04-14T18:53:27.927084Z","iopub.status.idle":"2024-04-14T18:53:28.040526Z","shell.execute_reply.started":"2024-04-14T18:53:27.927051Z","shell.execute_reply":"2024-04-14T18:53:28.039619Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Transfer the val images\n\nfor image in val_list:\n    \n    # the id in the csv file does not have the .tif extension therefore we add it here\n    fname = image + '.tif'\n    # get the label for a certain image\n    target = df_sam.loc[image,'label']\n    \n    # these must match the folder names\n    if target == 0:\n        label = 'a_no_tumor_tissue'\n    if target == 1:\n        label = 'b_has_tumor_tissue'\n    \n\n    # source path to image\n    src = os.path.join(list_l[3], fname)\n    # destination path to image\n    dst = os.path.join(dir_val, label, fname)\n    # copy the image from the source to the destination\n    shutil.copyfile(src, dst)","metadata":{"execution":{"iopub.status.busy":"2024-04-14T18:53:45.845313Z","iopub.execute_input":"2024-04-14T18:53:45.845678Z","iopub.status.idle":"2024-04-14T18:55:27.956298Z","shell.execute_reply.started":"2024-04-14T18:53:45.845651Z","shell.execute_reply":"2024-04-14T18:55:27.955212Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(len(os.listdir('base_dir/dir_val/a_no_tumor_tissue')))\nprint(len(os.listdir('base_dir/dir_val/b_has_tumor_tissue')))","metadata":{"execution":{"iopub.status.busy":"2024-04-14T18:57:40.723609Z","iopub.execute_input":"2024-04-14T18:57:40.723957Z","iopub.status.idle":"2024-04-14T18:57:40.739467Z","shell.execute_reply.started":"2024-04-14T18:57:40.723930Z","shell.execute_reply":"2024-04-14T18:57:40.738639Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Set up the generators\ntrain_path = 'base_dir/dir_tr'\nvalid_path = 'base_dir/dir_val'\ntest_path = list_l[2]\n\nnum_train_samples = len(df_sam_train)\nnum_val_samples = len(df_sam_val)\ntrain_batch_size = 10\nval_batch_size = 10\n\ntrain_steps = np.ceil(num_train_samples / train_batch_size)\nval_steps = np.ceil(num_val_samples / val_batch_size)","metadata":{"execution":{"iopub.status.busy":"2024-04-14T18:57:42.836252Z","iopub.execute_input":"2024-04-14T18:57:42.836602Z","iopub.status.idle":"2024-04-14T18:57:42.842084Z","shell.execute_reply.started":"2024-04-14T18:57:42.836575Z","shell.execute_reply":"2024-04-14T18:57:42.841054Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"datagen = ImageDataGenerator(rescale=1.0/255)\n\ntrain_gen = datagen.flow_from_directory(train_path,\n                                        target_size=(IMAGE_SIZE,IMAGE_SIZE),\n                                        batch_size=train_batch_size,\n                                        class_mode='categorical')\n\nval_gen = datagen.flow_from_directory(valid_path,\n                                        target_size=(IMAGE_SIZE,IMAGE_SIZE),\n                                        batch_size=val_batch_size,\n                                        class_mode='categorical')\n\n# Note: shuffle=False causes the test dataset to not be shuffled\ntest_gen = datagen.flow_from_directory(valid_path,\n                                        target_size=(IMAGE_SIZE,IMAGE_SIZE),\n                                        batch_size=1,\n                                        class_mode='categorical',\n                                        shuffle=False)","metadata":{"execution":{"iopub.status.busy":"2024-04-14T18:57:47.896930Z","iopub.execute_input":"2024-04-14T18:57:47.897299Z","iopub.status.idle":"2024-04-14T18:57:54.540666Z","shell.execute_reply.started":"2024-04-14T18:57:47.897270Z","shell.execute_reply":"2024-04-14T18:57:54.539923Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Model Configuration\n\n- **Kernel Size**: Each convolutional layer uses a kernel (filter) size of `(3,3)`, meaning the filters that convolve around the input have a dimension of 3x3 pixels.\n- **Pooling Size**: The max pooling layers use a pooling window of `(2,2)`, which will downsample the input representation by taking the maximum value over a 2x2 pooling window.\n- **Filters**: The number of filters in the convolutional layers starts with `32` and increases to `64` and then `128` as the network deepens. This is a common practice that helps the network learn increasingly complex features at each layer.\n- **Dropout**: Dropout layers are added with a dropout rate of `0.3`, which means during training, 30% of the units in the layer's output are randomly set to zero. This helps prevent overfitting by ensuring that the network does not rely on any one feature.\n\n### Sequential Model Layers\n\n1. **Input Layer**: The input shape is specified as `(96, 96, 3)`, suitable for RGB images of size 96x96 pixels.\n2. **Convolutional and Pooling Layers**: The model consists of three blocks of three convolutional layers each, followed by a max pooling layer. Each convolutional layer uses `ReLU` (Rectified Linear Unit) as the activation function.\n3. **Flattening**: After the convolutional and pooling layers, a `Flatten` layer is used to convert the 2D feature maps to a 1D vector, making it possible to feed the features into the dense layers.\n4. **Dense Layers**: A dense layer with `256` units and `ReLU` activation function is used, followed by a dropout layer.\n5. **Output Layer**: The final dense layer has `2` units with a `softmax` activation function, which is appropriate for binary classification tasks. It outputs the probability distribution over the two classes.\n\n### Model Summary\nAfter defining the model, `model.summary()` is called to print the details of the model architecture, including the number of parameters and the shape of the output at each layer.\n\n```python\nmodel = Sequential()\nmodel.add(Conv2D(first_filters, kernel_size, activation = 'relu', input_shape = (96, 96, 3)))\n...\nmodel.add(Dense(2, activation = \"softmax\"))\n\nmodel.summary()\n","metadata":{}},{"cell_type":"code","source":"kernel_size = (3,3)\npool_size= (2,2)\nfirst_filters = 32\nsecond_filters = 64\nthird_filters = 128\n\ndropout_conv = 0.3\ndropout_dense = 0.3\n\n\nmodel = Sequential()\nmodel.add(Conv2D(first_filters, kernel_size, activation = 'relu', input_shape = (96, 96, 3)))\nmodel.add(Conv2D(first_filters, kernel_size, activation = 'relu'))\nmodel.add(Conv2D(first_filters, kernel_size, activation = 'relu'))\nmodel.add(MaxPooling2D(pool_size = pool_size)) \nmodel.add(Dropout(dropout_conv))\n\nmodel.add(Conv2D(second_filters, kernel_size, activation ='relu'))\nmodel.add(Conv2D(second_filters, kernel_size, activation ='relu'))\nmodel.add(Conv2D(second_filters, kernel_size, activation ='relu'))\nmodel.add(MaxPooling2D(pool_size = pool_size))\nmodel.add(Dropout(dropout_conv))\n\nmodel.add(Conv2D(third_filters, kernel_size, activation ='relu'))\nmodel.add(Conv2D(third_filters, kernel_size, activation ='relu'))\nmodel.add(Conv2D(third_filters, kernel_size, activation ='relu'))\nmodel.add(MaxPooling2D(pool_size = pool_size))\nmodel.add(Dropout(dropout_conv))\n\nmodel.add(Flatten())\nmodel.add(Dense(256, activation = \"relu\"))\nmodel.add(Dropout(dropout_dense))\nmodel.add(Dense(2, activation = \"softmax\"))\n\nmodel.summary()","metadata":{"execution":{"iopub.status.busy":"2024-04-14T18:58:19.660813Z","iopub.execute_input":"2024-04-14T18:58:19.661537Z","iopub.status.idle":"2024-04-14T18:58:20.808165Z","shell.execute_reply.started":"2024-04-14T18:58:19.661506Z","shell.execute_reply":"2024-04-14T18:58:20.807276Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"model.compile(Adam(learning_rate=0.0001), loss='binary_crossentropy', metrics=['accuracy'])\n","metadata":{"execution":{"iopub.status.busy":"2024-04-14T18:58:25.176930Z","iopub.execute_input":"2024-04-14T18:58:25.178051Z","iopub.status.idle":"2024-04-14T18:58:25.192847Z","shell.execute_reply.started":"2024-04-14T18:58:25.178009Z","shell.execute_reply":"2024-04-14T18:58:25.192015Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"filepath = \"model.keras\"\ncheckpoint = ModelCheckpoint(filepath, monitor='val_acc', verbose=1, \n                             save_best_only=True, mode='max')\n\nreduce_lr = ReduceLROnPlateau(monitor='val_acc', factor=0.5, patience=2, \n                                   verbose=1, mode='max', min_lr=0.00001)\n                              \n                              \ncallbacks_list = [checkpoint, reduce_lr]\n\nhistory = model.fit(\n    train_gen, steps_per_epoch=int(train_steps),  \n    validation_data=val_gen, validation_steps=int(val_steps),  \n    epochs=20, verbose=1, callbacks=callbacks_list)","metadata":{"execution":{"iopub.status.busy":"2024-04-14T19:10:32.391044Z","iopub.execute_input":"2024-04-14T19:10:32.391431Z","iopub.status.idle":"2024-04-14T19:36:54.565555Z","shell.execute_reply.started":"2024-04-14T19:10:32.391403Z","shell.execute_reply":"2024-04-14T19:36:54.564760Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"val_loss, val_acc = model.evaluate(test_gen, steps=len(df_sam_val))\n\nprint('val_loss:', val_loss)\nprint('val_acc:', val_acc)","metadata":{"execution":{"iopub.status.busy":"2024-04-14T19:37:03.464592Z","iopub.execute_input":"2024-04-14T19:37:03.464946Z","iopub.status.idle":"2024-04-14T19:37:39.902267Z","shell.execute_reply.started":"2024-04-14T19:37:03.464918Z","shell.execute_reply":"2024-04-14T19:37:39.901390Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"acc = [x for x in history.history['accuracy'] if x > 0]\nval_acc = [x for x in history.history['val_accuracy'] if x > 0]\nloss = [x for x in history.history['loss'] if x > 0]\nval_loss = [x for x in history.history['val_loss'] if x > 0]\n\nepochs = range(1, len(acc) + 1)\n\nplt.plot(epochs, loss, 'bo', label='Training loss')\nplt.plot(epochs, val_loss, 'b', label='Validation loss')\nplt.title('Training and validation loss')\nplt.legend()\nplt.figure()\n\nplt.plot(epochs, acc, 'bo', label='Training acc')\nplt.plot(epochs, val_acc, 'b', label='Validation acc')\nplt.title('Training and validation accuracy')\nplt.legend()\nplt.figure()","metadata":{"execution":{"iopub.status.busy":"2024-04-14T19:37:46.406692Z","iopub.execute_input":"2024-04-14T19:37:46.407057Z","iopub.status.idle":"2024-04-14T19:37:46.921421Z","shell.execute_reply.started":"2024-04-14T19:37:46.407029Z","shell.execute_reply":"2024-04-14T19:37:46.920475Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Validation","metadata":{}},{"cell_type":"code","source":"predictions = model.predict(test_gen, steps=len(df_sam_val), verbose=1)\npredictions.shape","metadata":{"execution":{"iopub.status.busy":"2024-04-14T19:37:55.564957Z","iopub.execute_input":"2024-04-14T19:37:55.565343Z","iopub.status.idle":"2024-04-14T19:38:38.557337Z","shell.execute_reply.started":"2024-04-14T19:37:55.565313Z","shell.execute_reply":"2024-04-14T19:38:38.556385Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_preds = pd.DataFrame(predictions, columns=['no_tumor_tissue', 'has_tumor_tissue'])\ndf_preds.head()","metadata":{"execution":{"iopub.status.busy":"2024-04-14T19:38:38.558873Z","iopub.execute_input":"2024-04-14T19:38:38.559172Z","iopub.status.idle":"2024-04-14T19:38:38.570072Z","shell.execute_reply.started":"2024-04-14T19:38:38.559147Z","shell.execute_reply":"2024-04-14T19:38:38.569124Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_dir = 'test_dir'\nos.mkdir(test_dir)\n    \n# create test_images inside test_dir\ntest_images = os.path.join(test_dir, 'test_images')\nos.mkdir(test_images)\n# check that the directory we created exists\nos.listdir('test_dir')","metadata":{"execution":{"iopub.status.busy":"2024-04-14T19:42:13.703464Z","iopub.execute_input":"2024-04-14T19:42:13.704745Z","iopub.status.idle":"2024-04-14T19:42:13.712330Z","shell.execute_reply.started":"2024-04-14T19:42:13.704700Z","shell.execute_reply":"2024-04-14T19:42:13.711499Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_list = os.listdir(test)\nfor image in test_list:    \n    fname = image    \n    src = os.path.join('../input/histopathologic-cancer-detection/test', fname)\n    dst = os.path.join(test_images, fname)\n    shutil.copyfile(src, dst)\n\nlen(os.listdir('test_dir/test_images'))","metadata":{"execution":{"iopub.status.busy":"2024-04-14T19:42:17.702949Z","iopub.execute_input":"2024-04-14T19:42:17.703911Z","iopub.status.idle":"2024-04-14T19:47:54.886201Z","shell.execute_reply.started":"2024-04-14T19:42:17.703876Z","shell.execute_reply":"2024-04-14T19:47:54.885303Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_gen = datagen.flow_from_directory('test_dir',target_size=(IMAGE_SIZE,IMAGE_SIZE),\n                                        batch_size=1,\n                                        class_mode='categorical',\n                                        shuffle=False)","metadata":{"execution":{"iopub.status.busy":"2024-04-14T19:48:29.339029Z","iopub.execute_input":"2024-04-14T19:48:29.339715Z","iopub.status.idle":"2024-04-14T19:48:31.122663Z","shell.execute_reply.started":"2024-04-14T19:48:29.339684Z","shell.execute_reply":"2024-04-14T19:48:31.121880Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"predictions = model.predict(test_gen, steps=len(os.listdir('test_dir/test_images')), verbose=1)","metadata":{"execution":{"iopub.status.busy":"2024-04-14T19:48:33.438087Z","iopub.execute_input":"2024-04-14T19:48:33.439153Z","iopub.status.idle":"2024-04-14T19:51:33.304169Z","shell.execute_reply.started":"2024-04-14T19:48:33.439117Z","shell.execute_reply":"2024-04-14T19:51:33.303279Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_preds = pd.DataFrame(predictions, columns=['no_tumor_tissue', 'has_tumor_tissue'])\ndf_preds['file_names']=test_gen.filenames\ndf_preds['id'] = df_preds['file_names'].str.split('/').str[-1].str.split('.').str[0]\n\ndf_preds.head()\nsubmission = pd.DataFrame({'id':df_preds['id'], \n                           'label':df_preds['has_tumor_tissue'], \n                          }).set_index('id')\n\nsubmission.to_csv('submission.csv', columns=['label']) \nsubmission.head()","metadata":{"execution":{"iopub.status.busy":"2024-04-14T19:58:49.756274Z","iopub.execute_input":"2024-04-14T19:58:49.757093Z","iopub.status.idle":"2024-04-14T19:58:50.385998Z","shell.execute_reply.started":"2024-04-14T19:58:49.757064Z","shell.execute_reply":"2024-04-14T19:58:50.385038Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}