{"cells":[{"metadata":{"_uuid":"08c30283b7d0d2f1619f23ef6a145cf5ea9e3d98"},"cell_type":"markdown","source":"## 1. Pretrained Networks in Histopathology\n\nInspired by following papers\n\n* [ Classiﬁcation of Breast Cancer Histology Image using Ensemble of Pre-trained Neural Networks](https://link.springer.com/chapter/10.1007/978-3-319-93000-8_91)\n* [ Deep convolutional neural networks for computer-aided detection: CNN architectures, dataset characteristics and transfer learning](https://www.ncbi.nlm.nih.gov/pubmed/26886976)\n*  [Unregistered multiview mammogram analysis with pre-trained deep learning models](https://pdfs.semanticscholar.org/62dd/e439c540e9cb8ed0d7d90fbbbb2340feceab.pdf)\n*  [Convolutional neural networks for medical image analysis: full training or ﬁne tuning?](https://arxiv.org/pdf/1706.00712.pdf)\n*  [Deep learning with non-medical training used for chest pathology identiﬁcation](https://www.cs.tau.ac.il/~wolf/papers/SPIE15chest.pdf)\n\nI would like to investigate pre-trained net works for the following cancer detection task. The results of abovementioned papers show that  the application of  pre-trained CNNs with adequate ﬁne-tuning outperformed or, in the worst case, performed as well as CNNs designed from scratch. Also ﬁne-tuned CNNs are more robust to the size of training sets than CNNs trained from scratch. Training a CNN from a set of pre-trained weights is called ﬁne-tuning where a common practice is to replace the last fully connected layer of the pre-trained CNN with a new fully connected layer that has as many neurons as the number of classes in the new target application. In general, the early layers of a CNN learn low level image features, which are applicable to most vision tasks, but the late layers learn high level features. An effective ﬁne-tuning technique is to start from the last layer and then incrementally include more layers in the update process until the desired performance is reached. To avoid wrecking the learned weights of the convolutional layers, the classifier should be trained separatly. But I will skip this step and try to perform deep tuning of classifier and VGG16 layers at the same time as shown in the following figure:\n\n<img src=\"https://i.ibb.co/sHLttcx/VGG16-4.jpg\" alt=\"VGG16-4\" border=\"0\" />"},{"metadata":{"_uuid":"e31fec82c12303d4119490bc0fadd2e8a26263ee"},"cell_type":"markdown","source":"## 2. Data Preprocessing\n- Dataframe containing all image identification names and corresponding labels is created\n- Two error images are removed\n- True positive and true negative diagnosis cases are created\n- The imbalanced data problem is solved by undersampling\n- Data set is splitted to 90% of train sets and 10% of validation set\n- The Image Generator method called flow_from_directory is used \n- The directory structure for binary classification problem is created\n- The train and validation sets are splitted to true and negative classes and moved to directories for flow_from_directory method"},{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true},"cell_type":"code","source":"import pandas as pd\nimport os\n\n# Check folder\nprint(os.listdir(\"../input\"))\n# print(os.listdir(\"../working/main/train/true_positive\"))\n\n# Save train labels to dataframe\ndf = pd.read_csv(\"../input/histopathologic-cancer-detection/train_labels.csv\")\n\n# Remove error image\ndf = df[df['id'] != 'dd6dfed324f9fcb6f93f46f32fc800f2ec196be2']\n\n# Remove error black image\ndf = df[df['id'] != '9369c7278ec8bcc6c880d99194de09fc2bd4efbe']\n\n# Show first rows\ndf.head()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"3b27e40809664936a1fc30cd629ba7598bb022db"},"cell_type":"markdown","source":"### 2.1. True Positive & True Negative Diagnosis\n- Define true positive cases (= has cancer) \n- Define tue negative cases (= no cancer) "},{"metadata":{"trusted":true,"_uuid":"89515c262b9716ead7f4902e924e4744d7d8130a"},"cell_type":"code","source":"# True positive diagnosis\ndf_true_positive = df[df[\"label\"] == 1]\nprint(\"True positive diagnosis: \" + str(len(df_true_positive)))\n\n# True negative diagnosis\ndf_true_negative = df[df[\"label\"] == 0]\nprint(\"True negative diagnosis: \" + str(len(df_true_negative)))","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"7e76ea6ffd0b113d08c8f5f8cfb224ade51b47f9"},"cell_type":"markdown","source":"###  2.2. Imbalanced Data\n\nAs you can see, there is some imbalance in the dataset (1.46:1) where true positive class is under-represented. Empirical  studies as presented in:\n*  [An improved algorithm for neural network classification of imbalanced training sets](https://ieeexplore.ieee.org/document/286891)\n\non training of Neural Networks with imbalanced data show that the computed net error gradient vector is dominated by the bigger class.\n\nThe most common solution at the data level includes many different appoaches of resampling.  Simplest methods as presented in:\n*  [An overview of classification algorithms for imbalanced datasets](https://pdfs.semanticscholar.org/239b/2210b3fbc1f4b8246437a88a668bf9a0d2c0.pdf)\n\nare called oversampling/undersampling where the size of the minority/dominant class is increased/reduced by randomly replications/deleting. Due to kernel limits I decided to use undersampling method.\n\n### 2.3. Undersampling"},{"metadata":{"trusted":true,"_uuid":"16f521f21a4a5fa6a0c0d8351c9e16143dd75651"},"cell_type":"code","source":"# Undersampling of dominant class to reach a 1/1 balance\n\nfrom sklearn.utils import shuffle\n\ndf_true_negative = shuffle(df_true_negative)\ndf_true_negative = df_true_negative[0:88800]\n\ndf_true_positive = shuffle(df_true_positive)\ndf_true_positive = df_true_positive[0:88800]\n\n# concat the dataframes\ndf = pd.concat([df_true_negative, df_true_positive], axis=0)\n\n# shuffle\ndf = shuffle(df)\ndf['label'].value_counts()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"b9f25429d5b1c57d558a262f0ff1f7f00957a412"},"cell_type":"markdown","source":"### 2.4. Split data to train and validation sets\n"},{"metadata":{"trusted":true,"_uuid":"c48e2b3087f144aa9dfe2e639f1ff95985e4d602"},"cell_type":"code","source":"# Split data set  to train and validation sets\nfrom sklearn.model_selection import train_test_split\n\n# Use stratify= df['label'] to get balance ratio 1/1 in train and validation sets\ndf_train, df_val = train_test_split(df, test_size=0.2, stratify= df['label'])\n\n# Check balancing\nprint(\"True positive in train data: \" +  str(len(df_train[df_train[\"label\"] == 1])))\nprint(\"True negative in train data: \" +  str(len(df_train[df_train[\"label\"] == 0])))\nprint(\"True positive in validation data: \" +  str(len(df_val[df_val[\"label\"] == 1])))\nprint(\"True negative in validation data: \" +  str(len(df_val[df_val[\"label\"] == 0])))","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"5c9d6e5dba069c86b92a9dbbe4be51460b850902"},"cell_type":"markdown","source":"### 2.5. Data Flow from Directory\nTo avoid a crashing of the kaggle kernel due to the RAM limits it is highly recommended to use the flow_from_directory() method of ImageDataGenerator class as in details presented in\n\n* [CNN - How to use 160,000 images without crashing](https://www.kaggle.com/vbookshelf/cnn-how-to-use-160-000-images-without-crashing)\n\n* [Tutorial on using Keras flow_from_directory and generators](https://medium.com/@vijayabhaskar96/tutorial-image-classification-with-keras-flow-from-directory-and-generators-95f75ebe5720)\n\nIt is neccerary to prepare directory structure according to this image:\n\n<img src=\"https://i.ibb.co/VVhmQmF/Directory.jpg\" alt=\"Directory\" border=\"0\">\n\n### 2.6. Directory Structure"},{"metadata":{"trusted":true,"_uuid":"aace1235b525fc42d97034c87ee70042af2b5367"},"cell_type":"code","source":"# Delete directory\nimport shutil\nshutil.rmtree('main', ignore_errors=True)\n\n# Create directory\nos.mkdir('main')\n\n# Create subfolder for train and val images\nos.mkdir(os.path.join('main', 'train'))\nos.mkdir(os.path.join('main', 'val'))\n\n# Create subfolders for true positive and true negative in train\nos.mkdir(os.path.join('main','train','true_positive'))\nos.mkdir(os.path.join('main','train','true_negative'))      \n         \n# Create subfolders for true positive and true negative in val\nos.mkdir(os.path.join('main','val','true_positive'))\nos.mkdir(os.path.join('main','val','true_negative'))","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"6ea3968a821b23c12cf81b2a0b01dbeb56b5637b"},"cell_type":"markdown","source":"### 2.7.  Image Classes for Directory Structure"},{"metadata":{"trusted":true,"_uuid":"d9e3960b83fd69b432f66881b0fb56869e461ef8"},"cell_type":"code","source":"# Prepare image name classes for the directory structure\n# Save all train true positive names to list and add .tif\ntrain_true_positive = df_train[df_train[\"label\"] == 1]['id'].tolist()\ntrain_true_positive = [name + \".tif\" for name in train_true_positive]\n\n# Save all train true negativeto names list and add .tif\ntrain_true_negative = df_train[df_train[\"label\"] == 0]['id'].tolist()\ntrain_true_negative = [name + \".tif\" for name in train_true_negative]\n\n# Save all val true positive \"id\" to list and add .tif\nval_true_positive = df_val[df_val[\"label\"] == 1]['id'].tolist()\nval_true_positive = [name + \".tif\" for name in val_true_positive]\n\n# Save all val true negative \"id\" to list\nval_true_negative = df_val[df_val[\"label\"] == 0]['id'].tolist()\nval_true_negative = [name + \".tif\" for name in val_true_negative]","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"4e1de174133e70ad385531ad1d52da581f974f77"},"cell_type":"markdown","source":"### 2.8. Random Visualization of True Positive and True Negative Cases"},{"metadata":{"trusted":true,"_uuid":"6848ed579cf421c15a0b98125f0b0bd7b138d1dd"},"cell_type":"code","source":"from PIL import Image\nfrom PIL import ImageDraw\nfrom sklearn.utils import shuffle\nimport matplotlib.pyplot as plt\n\nimage_numbers = 3 \nfig = plt.figure()\nfig, ax = plt.subplots(2,image_numbers, figsize=(18,8))\n\nfor i in range(0,image_numbers):\n    \n    # Random images with true positive diagnosis\n    pos_train_names = shuffle(val_true_positive)\n    image = Image.open(os.path.join(\"../input/histopathologic-cancer-detection/train\",pos_train_names[i]))\n    print(image.size)\n    draw = ImageDraw.Draw(image)\n    draw.rectangle([(32,32),(64,64)], outline=\"yellow\")\n    ax[0,i].imshow(image)\n    ax[0,i].set_title(\"True Positive\",fontsize=14)\n    \n    # Random images with true negative diagnosis\n    neg_train_names = shuffle(train_true_negative)\n    image = Image.open(os.path.join(\"../input/histopathologic-cancer-detection/train\",neg_train_names[i]))\n    draw = ImageDraw.Draw(image)\n    draw.rectangle([(32,32),(64,64)], outline=\"yellow\")\n    ax[1,i].imshow(image)\n    ax[1,i].set_title(\"True Negative\",fontsize=14)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"ea8649b1bfcf8621a51c7eb8ad4fd46c3afc1221"},"cell_type":"markdown","source":"### 2.9. Moving images to directory folders"},{"metadata":{"trusted":true,"_uuid":"1be3ab8598f83bd8578d00a05f2e5318c7397f04"},"cell_type":"code","source":"# Move images to directory structure\nimport shutil\nimport os\nfrom tqdm import tqdm\n\ndef transfer(source,destination,files):\n    for image in tqdm(files):\n        # source path to image\n        src = os.path.join(source,image)\n        dst = os.path.join(destination,image)\n        # copy the image from the source to the destination\n        shutil.copyfile(src,dst)\n        \n# transfer\ntransfer('../input/histopathologic-cancer-detection/train','main/train/true_positive',train_true_positive)\ntransfer('../input/histopathologic-cancer-detection/train','main/train/true_negative',train_true_negative)\ntransfer('../input/histopathologic-cancer-detection/train','main/val/true_positive',val_true_positive)\ntransfer('../input/histopathologic-cancer-detection/train','main/val/true_negative',val_true_negative)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"3bcc7f9601367b9602f6ae4ed193a3623cfccf10"},"cell_type":"markdown","source":"## 3. Processing\n### 3.1. Fine Tuning of VGG16 Model\n\n<img src=\"https://i.ibb.co/sHLttcx/VGG16-4.jpg\" alt=\"VGG16-4\" border=\"0\" />\n\n### Load VGG16"},{"metadata":{"trusted":true,"_uuid":"ecd35ec4836d14e0aa9d8791f9329fc883de1caa"},"cell_type":"code","source":"# Import VGG16 model, with weights pre-trained on ImageNet.\nfrom keras.applications.vgg16 import VGG16, preprocess_input\n\n# VGG model without the last classifier layers (include_top = False)\nvgg16_model = VGG16(include_top = False,\n                    input_shape = (96,96,3),\n                    weights='../input/vgg16/vgg16_weights_tf_dim_ordering_tf_kernels_notop.h5')\n    \n# Freeze the layers \nfor layer in vgg16_model.layers[:-12]:\n    layer.trainable = False\n    \n# Check the trainable status of the individual layers\nfor layer in vgg16_model.layers:\n    print(layer, layer.trainable)\n    ","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"9d0f787a2ca9b6151d9be5223a7f3893ecf68bfb"},"cell_type":"markdown","source":"### 3.2. Define New Classifier"},{"metadata":{"trusted":true,"_uuid":"63111e2859611d21381d3d563d978dc6478b07ad"},"cell_type":"code","source":"from keras.models import Sequential\nfrom keras.layers import Dense,Flatten,Dropout\n\nmodel = Sequential()\nmodel.add(vgg16_model)\nmodel.add(Flatten())\nmodel.add(Dense(1024, activation=\"relu\"))\nmodel.add(Dropout(0.5))\nmodel.add(Dense(512, activation=\"relu\"))\nmodel.add(Dropout(0.5))\nmodel.add(Dense(2, activation=\"softmax\"))\nmodel.summary()\n# Load previous weights to save the time\n# model.load_weights('../input/pretrained3/model.h5')","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"61f2d445b39364240309c689fb1315fe8531ba50"},"cell_type":"markdown","source":"### 3.3 Data Generators and Augmentation\nAccording to the paper [Data Augmentation Techniques for Medical Imaging Classification Tasks](https://www.ncbi.nlm.nih.gov/pmc/articles/PMC5977656/) shear, rotate, scale, flips, jitter, gaussian filter augmentation sets correlate to a higher classification accuracy while noice and powers had much lower accuracies. "},{"metadata":{"trusted":true,"_uuid":"8a2075dcad7ba98e7d3d51945b253d4b37ef1a57"},"cell_type":"code","source":"# Generate batches of tensor image data with real-time data augmentation. \nimport numpy as np\nnum_train_samples = len(df_train)\nnum_val_samples = len(df_val)\ntrain_batch_size = 32\nval_batch_size = 32\n\ntrain_steps = np.ceil(num_train_samples / train_batch_size)\nval_steps = np.ceil(num_val_samples / val_batch_size)\n\nprint(train_steps)\nprint(val_steps)\n\nfrom tensorflow.keras.preprocessing.image import ImageDataGenerator\n\n# Augmentation \ntrain_datagen = ImageDataGenerator(\n                rescale=1./255,\n                vertical_flip=True,\n                horizontal_flip=True,\n                rotation_range=90,\n                shear_range=0.05)\n\n# Augmentation \n# train_datagen = ImageDataGenerator(\n#                 rescale=1./255,\n#                 vertical_flip=True, # set True or False\n#                 horizontal_flip=True, # set True or False\n#                 rotation_range=180, # 0-180 range, degrees of rotatation\n#                 shear_range=0.01,# 0-1 range for shearing\n#                 width_shift_range=0.2, # 0-1 range, horizontal translation\n#                 height_shift_range=0.2) # 0-1 range vertical translation\n\n#train_datagen = ImageDataGenerator(rescale=1./255)\n\n# Augmentation configuration for validatiopn and testing: only rescaling!\ntest_datagen = ImageDataGenerator(rescale=1./255)\n\n# Generator that will read pictures found in subfolers of 'main/train', and indefinitely generate batches of augmented image data\ntrain_generator = train_datagen.flow_from_directory('main/train',\n                                            target_size=(96,96),\n                                            batch_size=train_batch_size,\n                                            class_mode='categorical')\n\nval_generator = test_datagen.flow_from_directory('main/val',\n                                            target_size=(96,96),\n                                            batch_size=val_batch_size,\n                                            class_mode='categorical')\n\n# !!! batch_size=1 & shuffle=False !!!!\ntest_generator = test_datagen.flow_from_directory('main/val',\n                                            target_size=(96,96),\n                                            batch_size=1,\n                                            class_mode='categorical',\n                                            shuffle=False)\n\n# Get the labels that are associated with each index\nprint(val_generator.class_indices)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"1db8162fbbbf096829f5d9c5ab7031cf6820c6f4"},"cell_type":"markdown","source":"### 3.4. Train model"},{"metadata":{"trusted":true,"_uuid":"b5337c54f5ef7fbd8046a37b07545d2035bd42b2","scrolled":true},"cell_type":"code","source":"from keras.optimizers import Adam, SGD\nfrom keras import optimizers\n\n#model.compile(Adam(lr=0.00001), loss='binary_crossentropy', metrics=['acc'])\nmodel.compile(loss='binary_crossentropy',optimizer=optimizers.SGD(lr=0.00001, momentum=0.95),metrics=['accuracy'])\n\n# Due to Disk limits the saving of best weights isn't possible during the training process                                            \nhistory = model.fit_generator(\n                    train_generator, \n                    steps_per_epoch  = train_steps, \n                    validation_data  = val_generator,\n                    validation_steps = val_steps,\n                    epochs           = 1, \n                    verbose          = 1)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"c68804969e5afd1973304c5842df61854b36cf88"},"cell_type":"markdown","source":"### 3.5. Monitoring of Validation Accuracy"},{"metadata":{"trusted":true,"_uuid":"4db354f3fad1a387e9b679742d79007e6da9c55e"},"cell_type":"code","source":"# Plot validation and accuracies over epochs\nimport matplotlib.pyplot as plt\n\ntrain_acc = history.history['acc']\nval_acc = history.history['val_acc']\n\nepochs = range(len(train_acc))\n\nplt.plot(epochs,train_acc,'b',label='Training accuracy')\nplt.plot(epochs,val_acc,'r',label='Validation accuracy')\nplt.title('Training and validation accuracy')\nplt.legend()\n\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"ca7efb78d06b9503b7884fc6725e2f9f6bad39be"},"cell_type":"code","source":"# Due to the disk limits I couldn't save the best model during the training process\nprint(\"Validation Accuracy: \" + str(history.history['val_acc'][-1:]))","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"f37dbea2fe05f5d73d2587ea1e1b455257b74769"},"cell_type":"markdown","source":"## 4. Postprocessing\n### 4.1. Confusion Matrix\nFor classification with the class imbalance problem, validation accuracy is no longer a proper measure since the under-represented class has less impact on the accuracy [[paper]](https://www.researchgate.net/publication/263913891_Classification_of_imbalanced_data_a_review ). In the bi-class scenario, samples can be categorized into four groups after a classification process:\n\n1. True Positive: In reality, the patient has cancer and the trained model classifies it correctly.\n\n2. False Negative: In reality, the patient has cancer but the model  doesn't classify it correctly.\n\n3. False Positive: In reality the patient is healthy but the model doesn't classify it correctly.\n\n3. True Negative: In reality the patient is healthy and the trained model classifies it correctly.\n\nAs you can see, the 2. scenario is the worst case in the context of the cancer detection problem. These groups can be denoted in the confusion matrix as shown in following code:"},{"metadata":{"trusted":true,"_uuid":"947a10dd545f34bec69cc7cb52d7f7b62b5a0549"},"cell_type":"code","source":"# Prediction on validation data sets\nval_predict = model.predict_generator(test_generator, steps=len(df_val), verbose=1)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"2b98a8a8fdec21c2ee2c26d96d7b9b336e41d22b","scrolled":true},"cell_type":"code","source":"from sklearn.metrics import confusion_matrix\nprint(\"Confusion matrix is: \")\nprint(confusion_matrix(test_generator.classes, val_predict.argmax(axis=1)))","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"c670c8d6f4dcc205e2d8521b6f6ea3a0debd2611"},"cell_type":"markdown","source":"Format\n\n|Model: negative prediction | Model: positive prediction\n--------------------|-\n**Reality: not true** | True negative TN| False positive FP\n**Reality: true** | False negative FN| True positive TP"},{"metadata":{"_uuid":"b3deae1b2b19df531f26e6f50bb07799c0b7e56f"},"cell_type":"markdown","source":"### 4.2. ROC AUC Analysis\nIn medical image classification the area under the ROC curve (AUC) is usually used as a performance metric for the goodness of an ROC curve. The AUC varies between 0 and 1. The greater the AUC, the greater the average probability of making a correct diagnosis."},{"metadata":{"trusted":true,"_uuid":"0a2a4dfb91caf513664e9e0dcc0ac6b0b5aa0e4c"},"cell_type":"code","source":"from sklearn.metrics import roc_curve, auc\n\nfpr, tpr, thresholds = roc_curve(test_generator.classes, val_predict.argmax(axis=1))   \n# Compute ROC area\nprint(\"ROC area is: \" + str(auc(fpr, tpr)))\n\nplt.figure()\nplt.plot(fpr, tpr, color='darkred', label='ROC curve (area = %0.2f)' % auc(fpr, tpr))\nplt.plot([0, 1], [0, 1], color='darkblue', linestyle='--')\nplt.xlim([-0.01, 1.0])\nplt.ylim([0.0, 1.01])\nplt.xlabel('False Positive Rate')\nplt.ylabel('True Positive Rate')\nplt.title('Receiver Operating Characteristic')\nplt.legend()\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"bf341c30bbc9da192cdb74810ebe94b46b31e437"},"cell_type":"markdown","source":"### 4.3. Test Set Prediction\n"},{"metadata":{"trusted":true,"_uuid":"e03bdf513624b1b216046faea8e42bb57068e7c3"},"cell_type":"code","source":"from tqdm import tqdm\n\n# Delete directory structure\nshutil.rmtree('main')\n\n# Create test folder\nos.mkdir('test')\n    \n# Create test_images folder inside test folder\nos.mkdir(os.path.join('test','test_images'))\n\n# Save test image identification names\ntest_images = os.listdir('../input/histopathologic-cancer-detection/test')\n\n# Move images to test folder\nfor test_image in tqdm(test_images):   \n    # source \n    src = os.path.join('../input/histopathologic-cancer-detection/test',test_image)\n    # destination \n    dst = os.path.join('test/test_images',test_image)\n    # copy the image\n    shutil.copyfile(src, dst)\n    \n# !!! batch_size=1 & shuffle=False !!!!\ntest_generator = test_datagen.flow_from_directory('test',\n                                            target_size=(96,96),\n                                            batch_size=1,\n                                            class_mode='categorical',\n                                            shuffle=False)\n\n# model.load_weights('best_model.h5')\n# Predict \ntest_predict = model.predict_generator(test_generator, steps=57458, verbose=1)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"8668d2a8f70c47b90f491489117a72149504d40f"},"cell_type":"markdown","source":" ### 4.4. Submission\n \n According to the rules, for each id in the test set, a probability must be predicted that the center 32x32px region of a patch contains at least one pixel of tumor tissue. The file should contain a header and have the following format:\n \n id | label\n--------------------|-\n0b2ea2a822ad23fdb1b5dd26653da899fbd2c0d5| 0\n95596b92e5066c5c52466c90b69ff089b39f2737| 0\n "},{"metadata":{"trusted":true,"_uuid":"4ac9e310fbce3ea073534fb063c93dc98afcba4c"},"cell_type":"code","source":"# Extract test names from test_generator\ntest_names = [name.split(\"/\")[1].split(\".\")[0] for name in test_generator.filenames]\n\n# Create a dataframe\ndf_pred_test = pd.DataFrame(test_names, columns=[\"id\"])\n\n# !!! label == probability of true positive (has cancer) !!!\ndf_pred_test[\"label\"] = pd.DataFrame(test_predict[:,1])\n\ndf_pred_test.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"c48a711c8158c55cbfd77a89577340fcb274872c"},"cell_type":"code","source":"# Save sample submissions as dataframe\ndf_samples = pd.read_csv('../input/histopathologic-cancer-detection/sample_submission.csv')\n\n# Merge df_samples and df_pred_test using unique key combination through \"id\"\nsubmission = pd.merge(df_samples.drop(\"label\",axis=1), df_pred_test, on = 'id')\nsubmission.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"27efca87623cbbe8e5ed1e2dd1eaf6de6681d8c6"},"cell_type":"code","source":"shutil.rmtree('test')\n\n# Save submission file\nsubmission.to_csv(\"submission.csv\",index=False) \n\n# Due to the disk limits I can only save the last model\nmodel.save('model.h5')","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.6.6","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"}},"nbformat":4,"nbformat_minor":1}