{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Git hub with all code [there](https://github.com/Sonya-Shultz/DLcourse)","metadata":{}},{"cell_type":"markdown","source":"**Histopathologic Cancer Detection**\nIn this task we mast create CNN for binar clasification. We need identify metastatic cancer in small image patches taken from larger digital pathology scans.\n\n**About dataset**\nAll data we tacke from kaggle \"Histopathologic Cancer Detection\" competition. This dataset contains a large number of small (96x96px) pathology images to classify. And .csv file with file's names and linked to it class (0 or 1). We have train and test data in diferent folders.\n\n**What we need to do?**\nFor each id in the test set, We must predict a probability that center 32x32px region of a patch contains at least one pixel of tumor tissue. Answer format should be: id, label 0b2ea2a822ad23fdb1b5dd26653da899fbd2c0d5,0 95596b92e5066c5c52466c90b69ff089b39f2737,0 248e6738860e2ebcf6258cdc1f32f299e0c76914,0 etc.","metadata":{}},{"cell_type":"code","source":"import numpy as np\nimport pandas as pd\nimport os\nfrom matplotlib import pyplot as plt\nfrom matplotlib import image as mpimg\nfrom sklearn.model_selection import train_test_split\n\nimport tensorflow as tf\nfrom tensorflow.keras.models import Sequential\nfrom tensorflow import keras\nfrom tensorflow.keras.layers import * \nfrom tensorflow.keras.preprocessing.image import ImageDataGenerator","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2022-05-17T14:07:26.504553Z","iopub.execute_input":"2022-05-17T14:07:26.504824Z","iopub.status.idle":"2022-05-17T14:07:26.510748Z","shell.execute_reply.started":"2022-05-17T14:07:26.504795Z","shell.execute_reply":"2022-05-17T14:07:26.509373Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Read all data from train_label.csv and change id to file name.**","metadata":{}},{"cell_type":"code","source":"train_data_full = pd.read_csv('../input/histopathologic-cancer-detection/train_labels.csv', dtype=str)\nprint(\"Shape of dataset: \",train_data_full.shape)\n\ntrain_data_full.id = train_data_full.id + '.tif'\ntrain_data_full.head()","metadata":{"execution":{"iopub.status.busy":"2022-05-17T14:07:26.525215Z","iopub.execute_input":"2022-05-17T14:07:26.525406Z","iopub.status.idle":"2022-05-17T14:07:26.762521Z","shell.execute_reply.started":"2022-05-17T14:07:26.525383Z","shell.execute_reply":"2022-05-17T14:07:26.761657Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**For some statistics, lets see label distribution**","metadata":{}},{"cell_type":"code","source":"(train_data_full.label.value_counts() / len(train_data_full)).to_frame().sort_index().T","metadata":{"execution":{"iopub.status.busy":"2022-05-17T14:07:26.764191Z","iopub.execute_input":"2022-05-17T14:07:26.764516Z","iopub.status.idle":"2022-05-17T14:07:26.798275Z","shell.execute_reply.started":"2022-05-17T14:07:26.764477Z","shell.execute_reply":"2022-05-17T14:07:26.797133Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Lets see some data images example:**","metadata":{}},{"cell_type":"code","source":"print(\"Dataset image example\")\nh_path = '../input/histopathologic-cancer-detection/train'\nsample = train_data_full.sample(n=16).reset_index()\n\nplt.figure(figsize=(6,6))\n\nfor i, row in sample.iterrows():\n    img = mpimg.imread(f'../input/histopathologic-cancer-detection/train/{row.id}')\n    label = row.label\n    plt.subplot(4,4,i+1)\n    plt.imshow(img)\n    plt.text(0,-5,f'label {label}', color='k')\n\n    plt.axis('off')\n\nplt.tight_layout()\nplt.show()\n","metadata":{"execution":{"iopub.status.busy":"2022-05-17T14:07:26.800054Z","iopub.execute_input":"2022-05-17T14:07:26.800379Z","iopub.status.idle":"2022-05-17T14:07:27.534561Z","shell.execute_reply.started":"2022-05-17T14:07:26.800341Z","shell.execute_reply":"2022-05-17T14:07:27.533916Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Prepare data for training and validation.\n\nAlso create data loaders for training and validation.","metadata":{}},{"cell_type":"code","source":"# Divade into training and validation part\ntrain_data, valid_data = train_test_split(train_data_full, test_size=0.2, random_state=1, stratify=train_data_full.label)\n\n# Data loaders\ntrain_data_gen = ImageDataGenerator(rescale=1/255)\nvalidation_data_gen = ImageDataGenerator(rescale=1/255)\n\nBATCH_SIZE = 64\n\ntrain_loader = train_data_gen.flow_from_dataframe(\n    dataframe = train_data,\n    directory = h_path,\n    x_col = 'id',\n    y_col = 'label',\n    batch_size = BATCH_SIZE,\n    seed = 1,\n    shuffle = True,\n    class_mode = 'categorical',\n    target_size = (96,96)\n)\n\nvalid_loader = train_data_gen.flow_from_dataframe(\n    dataframe = valid_data,\n    directory = h_path,\n    x_col = 'id',\n    y_col = 'label',\n    batch_size = BATCH_SIZE,\n    seed = 1,\n    shuffle = True,\n    class_mode = 'categorical',\n    target_size = (96,96)\n)","metadata":{"execution":{"iopub.status.busy":"2022-05-17T14:07:27.536649Z","iopub.execute_input":"2022-05-17T14:07:27.537059Z","iopub.status.idle":"2022-05-17T14:12:58.832691Z","shell.execute_reply.started":"2022-05-17T14:07:27.537022Z","shell.execute_reply":"2022-05-17T14:12:58.831863Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# **Now build CNN.**\n(Set log level becouse of warning from kaggle)\n\n**Input shape** is 96px x 96px x 3(rgb)\n\n**Use 3 blocks like:**\n* two conv2d layers, \n* max pooling (2x2), \n* dropout and batch normalization. \n* Activation function is relu, padding - same.\n\n*Then flatten layer.*\n\n**Next for clasification:**\n* Dense 64 layer + dropout (activation - relu)\n* Dense 8 layer + dropout (activation - relu)\n* Batch normalization\n* Output layer - dense 2 with sigmoid activation (good for binary classification)","metadata":{}},{"cell_type":"code","source":"# Build CNN\nnp.random.seed(1)\ntf.random.set_seed(1)\n\nos.environ['TF_CPP_MIN_LOG_LEVEL'] = '3' \ncnn = Sequential([\n    Conv2D(64, (3,3), activation = 'relu', padding = 'same', input_shape=(96,96,3)),\n    Conv2D(64, (3,3), activation = 'relu', padding = 'same'),\n    MaxPooling2D(2,2),\n    Dropout(0.5),\n    BatchNormalization(),\n\n    Conv2D(64, (3,3), activation = 'relu', padding = 'same'),\n    Conv2D(64, (3,3), activation = 'relu', padding = 'same'),\n    MaxPooling2D(2,2),\n    Dropout(0.5),\n    BatchNormalization(),\n    \n    Conv2D(128, (3,3), activation = 'relu', padding = 'same'),\n    Conv2D(128, (3,3), activation = 'relu', padding = 'same'),\n    MaxPooling2D(2,2),\n    Dropout(0.5),\n    BatchNormalization(),\n\n    Flatten(),\n    \n    Dense(64, activation='relu'),\n    Dropout(0.5),\n    Dense(8, activation='relu'),\n    Dropout(0.5),\n    BatchNormalization(),\n    Dense(2, activation='sigmoid')\n])\n\ncnn.summary()","metadata":{"execution":{"iopub.status.busy":"2022-05-17T14:12:58.834128Z","iopub.execute_input":"2022-05-17T14:12:58.834572Z","iopub.status.idle":"2022-05-17T14:12:58.999282Z","shell.execute_reply.started":"2022-05-17T14:12:58.834531Z","shell.execute_reply":"2022-05-17T14:12:58.9979Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Set optimazer parametr for first training","metadata":{}},{"cell_type":"code","source":"opt = tf.keras.optimizers.Adam(0.001)\ncnn.compile(loss='categorical_crossentropy', optimizer=opt, metrics=['accuracy', tf.keras.metrics.AUC()])","metadata":{"execution":{"iopub.status.busy":"2022-05-17T14:12:59.000563Z","iopub.execute_input":"2022-05-17T14:12:59.000844Z","iopub.status.idle":"2022-05-17T14:12:59.015551Z","shell.execute_reply.started":"2022-05-17T14:12:59.00081Z","shell.execute_reply":"2022-05-17T14:12:59.014852Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Satrt training 40 epoch","metadata":{}},{"cell_type":"code","source":"%%time \n\nh1 = cnn.fit(\n    x = train_loader, \n    steps_per_epoch = len(train_loader), \n    epochs = 40,\n    validation_data = valid_loader, \n    validation_steps = len(valid_loader), \n    verbose = 1\n)","metadata":{"execution":{"iopub.status.busy":"2022-05-17T14:12:59.016657Z","iopub.execute_input":"2022-05-17T14:12:59.016984Z","iopub.status.idle":"2022-05-17T18:12:34.413247Z","shell.execute_reply.started":"2022-05-17T14:12:59.016942Z","shell.execute_reply":"2022-05-17T18:12:34.412491Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**See what we can show on plot**","metadata":{}},{"cell_type":"code","source":"history = h1.history\nprint(history.keys())","metadata":{"execution":{"iopub.status.busy":"2022-05-17T18:12:34.416279Z","iopub.execute_input":"2022-05-17T18:12:34.416475Z","iopub.status.idle":"2022-05-17T18:12:34.422292Z","shell.execute_reply.started":"2022-05-17T18:12:34.41645Z","shell.execute_reply":"2022-05-17T18:12:34.421469Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Function wich show us some training results","metadata":{}},{"cell_type":"code","source":"def show_res():\n    epoch_range = range(1, len(history['loss'])+1)\n\n    plt.figure(figsize=[14,4])\n    plt.subplot(1,3,1)\n    plt.plot(epoch_range, history['loss'], label='Training')\n    plt.plot(epoch_range, history['val_loss'], label='Validation')\n    plt.xlabel('Epoch'); plt.ylabel('Loss'); plt.title('Loss')\n    plt.legend()\n    plt.subplot(1,3,2)\n    plt.plot(epoch_range, history['accuracy'], label='Training')\n    plt.plot(epoch_range, history['val_accuracy'], label='Validation')\n    plt.xlabel('Epoch'); plt.ylabel('Accuracy'); plt.title('Accuracy')\n    plt.legend()\n    plt.subplot(1,3,3)\n    plt.plot(epoch_range, history['auc_1'], label='Training')\n    plt.plot(epoch_range, history['val_auc_1'], label='Validation')\n    plt.xlabel('Epoch'); plt.ylabel('AUC'); plt.title('AUC')\n    plt.legend()\n    plt.tight_layout()\n    plt.show()\n    \nshow_res()","metadata":{"execution":{"iopub.status.busy":"2022-05-17T18:16:49.412437Z","iopub.execute_input":"2022-05-17T18:16:49.413101Z","iopub.status.idle":"2022-05-17T18:16:49.948488Z","shell.execute_reply.started":"2022-05-17T18:16:49.413062Z","shell.execute_reply":"2022-05-17T18:16:49.947799Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Secound training\n**As see we had a \"problem\" with Validation set**\n\nWe, probably, have too big learning rate. So change it to \"normal\".\n\nFor not overfition our model let's do 20 epochs.","metadata":{}},{"cell_type":"code","source":"tf.keras.backend.set_value(cnn.optimizer.learning_rate, 0.0001)","metadata":{"execution":{"iopub.status.busy":"2022-05-17T18:18:43.082055Z","iopub.execute_input":"2022-05-17T18:18:43.082916Z","iopub.status.idle":"2022-05-17T18:18:43.088849Z","shell.execute_reply.started":"2022-05-17T18:18:43.082865Z","shell.execute_reply":"2022-05-17T18:18:43.087923Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time \n\nh2 = cnn.fit(\n    x = train_loader, \n    steps_per_epoch = len(train_loader), \n    epochs = 20,\n    validation_data = valid_loader, \n    validation_steps = len(valid_loader), \n    verbose = 1\n)","metadata":{"execution":{"iopub.status.busy":"2022-05-17T18:18:46.643361Z","iopub.execute_input":"2022-05-17T18:18:46.64391Z","iopub.status.idle":"2022-05-17T20:11:02.601491Z","shell.execute_reply.started":"2022-05-17T18:18:46.643873Z","shell.execute_reply":"2022-05-17T20:11:02.600541Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for k in history.keys():\n    history[k] += h2.history[k]\nshow_res()","metadata":{"execution":{"iopub.status.busy":"2022-05-17T20:11:08.789087Z","iopub.execute_input":"2022-05-17T20:11:08.789569Z","iopub.status.idle":"2022-05-17T20:11:09.279307Z","shell.execute_reply.started":"2022-05-17T20:11:08.789531Z","shell.execute_reply":"2022-05-17T20:11:09.278581Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Third training\n\n**We see that previous train give us much less loss and higher accuracy.**\n\nSo, let train it another time with same settings but for 10 epoch.","metadata":{}},{"cell_type":"code","source":"tf.keras.backend.set_value(cnn.optimizer.learning_rate, 0.0001)","metadata":{"execution":{"iopub.status.busy":"2022-05-17T20:13:11.629151Z","iopub.execute_input":"2022-05-17T20:13:11.629441Z","iopub.status.idle":"2022-05-17T20:13:11.636078Z","shell.execute_reply.started":"2022-05-17T20:13:11.629409Z","shell.execute_reply":"2022-05-17T20:13:11.635057Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time \n\nh3 = cnn.fit(\n    x = train_loader, \n    steps_per_epoch = len(train_loader), \n    epochs = 10,\n    validation_data = valid_loader, \n    validation_steps = len(valid_loader), \n    verbose = 1\n)","metadata":{"execution":{"iopub.status.busy":"2022-05-17T20:13:18.937603Z","iopub.execute_input":"2022-05-17T20:13:18.938079Z","iopub.status.idle":"2022-05-17T21:07:24.951104Z","shell.execute_reply.started":"2022-05-17T20:13:18.93804Z","shell.execute_reply":"2022-05-17T21:07:24.950337Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for k in history.keys():\n    history[k] += h3.history[k]\nshow_res()","metadata":{"execution":{"iopub.status.busy":"2022-05-17T21:07:39.818827Z","iopub.execute_input":"2022-05-17T21:07:39.819525Z","iopub.status.idle":"2022-05-17T21:07:40.592212Z","shell.execute_reply.started":"2022-05-17T21:07:39.819401Z","shell.execute_reply":"2022-05-17T21:07:40.591551Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**So, as we can see, accuracy preaty good (round 0.95 for validation set).**\n\nSo, let's stop on this.","metadata":{}},{"cell_type":"markdown","source":"# Submission\n**Firstly we need to read all data from test folder**\n\nAnd look some semple from it.\n\n*Not shuffle data becouse we goin predict not train with model.*","metadata":{}},{"cell_type":"code","source":"test_data = pd.read_csv('../input/histopathologic-cancer-detection/sample_submission.csv')\nprint('Test data has ', test_data.shape, ' size.')\ntest_data['file_n'] = test_data.id + '.tif'\ntest_data.head()","metadata":{"execution":{"iopub.status.busy":"2022-05-17T21:09:26.902952Z","iopub.execute_input":"2022-05-17T21:09:26.903202Z","iopub.status.idle":"2022-05-17T21:09:26.960113Z","shell.execute_reply.started":"2022-05-17T21:09:26.903175Z","shell.execute_reply":"2022-05-17T21:09:26.959276Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"BATCH_SIZE = 64\n\ntest_data_gen = ImageDataGenerator(rescale=1/255)\n\ntest_loader = test_data_gen.flow_from_dataframe(\n    dataframe = test_data,\n    directory = \"../input/histopathologic-cancer-detection/test\",\n    x_col = 'file_n',\n    batch_size = BATCH_SIZE,\n    shuffle = False,\n    class_mode = None,\n    target_size = (96,96)\n)","metadata":{"execution":{"iopub.status.busy":"2022-05-17T21:09:31.296252Z","iopub.execute_input":"2022-05-17T21:09:31.298884Z","iopub.status.idle":"2022-05-17T21:12:15.840889Z","shell.execute_reply.started":"2022-05-17T21:09:31.298839Z","shell.execute_reply":"2022-05-17T21:12:15.840123Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Start prediction and see size of result**","metadata":{}},{"cell_type":"code","source":"test_end = cnn.predict(test_loader)\nprint(test_end.shape)","metadata":{"execution":{"iopub.status.busy":"2022-05-17T21:12:22.874518Z","iopub.execute_input":"2022-05-17T21:12:22.875136Z","iopub.status.idle":"2022-05-17T21:19:32.806606Z","shell.execute_reply.started":"2022-05-17T21:12:22.875096Z","shell.execute_reply":"2022-05-17T21:19:32.805865Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"submission = pd.read_csv('../input/histopathologic-cancer-detection/sample_submission.csv')\nsubmission.head()","metadata":{"execution":{"iopub.status.busy":"2022-05-17T21:20:32.844418Z","iopub.execute_input":"2022-05-17T21:20:32.84485Z","iopub.status.idle":"2022-05-17T21:20:32.89555Z","shell.execute_reply.started":"2022-05-17T21:20:32.844811Z","shell.execute_reply":"2022-05-17T21:20:32.894647Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**lets fit our ansver to competition submission form**","metadata":{}},{"cell_type":"code","source":"submission.label = test_end[:,1]\nsubmission.head()","metadata":{"execution":{"iopub.status.busy":"2022-05-17T21:20:41.900187Z","iopub.execute_input":"2022-05-17T21:20:41.900437Z","iopub.status.idle":"2022-05-17T21:20:41.909943Z","shell.execute_reply.started":"2022-05-17T21:20:41.900407Z","shell.execute_reply":"2022-05-17T21:20:41.909231Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"submission.to_csv('submission.csv', header=True, index=False)","metadata":{"execution":{"iopub.status.busy":"2022-05-17T21:21:48.484494Z","iopub.execute_input":"2022-05-17T21:21:48.484766Z","iopub.status.idle":"2022-05-17T21:21:48.736231Z","shell.execute_reply.started":"2022-05-17T21:21:48.484732Z","shell.execute_reply":"2022-05-17T21:21:48.735314Z"},"trusted":true},"execution_count":null,"outputs":[]}]}