{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Histopathologic Cancer Detection with a Convolutional Neural Network\n\nIn this challenge, we are given >200,000 .tif files depicting (96 x 96) RGB microscope images of lymph tissue. We are tasked with building a model to determine whether malignant tissue is present in the centers of the images. This task is well-suited to convolutional neural networks, which are designed to be capable of internally representing interactions between structures consisting of regularly-spaced data.","metadata":{}},{"cell_type":"code","source":"import os\nfrom PIL import Image\nimport numpy as np\nimport pandas as pd\n\nfrom keras.preprocessing.image import ImageDataGenerator\nfrom keras.models import Sequential\nfrom keras.layers import Conv2D, BatchNormalization, Activation, MaxPooling2D, Dropout, Flatten, Dense\nfrom keras.regularizers import L2\nfrom keras.metrics import AUC","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2023-04-20T18:52:03.071704Z","iopub.execute_input":"2023-04-20T18:52:03.072208Z","iopub.status.idle":"2023-04-20T18:52:10.727135Z","shell.execute_reply.started":"2023-04-20T18:52:03.072159Z","shell.execute_reply":"2023-04-20T18:52:10.725962Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"top_dir = '/kaggle/input/histopathologic-cancer-detection/'\ntrain_dir = top_dir + 'train/'\ntest_dir = top_dir + 'test/'","metadata":{"execution":{"iopub.status.busy":"2023-04-20T18:52:10.729745Z","iopub.execute_input":"2023-04-20T18:52:10.730531Z","iopub.status.idle":"2023-04-20T18:52:10.735705Z","shell.execute_reply.started":"2023-04-20T18:52:10.730486Z","shell.execute_reply":"2023-04-20T18:52:10.734437Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def display_image(train_dir, ind):\n    test_image = os.listdir(train_dir)\n    filename = train_dir + test_image[ind]\n    return Image.open(filename)\n\ndisplay_image(train_dir, 0)","metadata":{"execution":{"iopub.status.busy":"2023-04-20T18:52:10.737632Z","iopub.execute_input":"2023-04-20T18:52:10.738394Z","iopub.status.idle":"2023-04-20T18:52:13.111493Z","shell.execute_reply.started":"2023-04-20T18:52:10.738351Z","shell.execute_reply":"2023-04-20T18:52:13.110522Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"y = pd.read_csv(top_dir + 'train_labels.csv')\ny['label'].hist()","metadata":{"execution":{"iopub.status.busy":"2023-04-20T18:52:13.113262Z","iopub.execute_input":"2023-04-20T18:52:13.114051Z","iopub.status.idle":"2023-04-20T18:52:13.774341Z","shell.execute_reply.started":"2023-04-20T18:52:13.114000Z","shell.execute_reply":"2023-04-20T18:52:13.773326Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Preliminary data analysis indicates that very little cleaning is necessary. Images with similar filenames do not appear to be correlated but will be shuffled regardless.\n\nThe data are reasonably balanced between positive and negative results.","metadata":{}},{"cell_type":"code","source":"# Construct data generators. Necessary due to memory limits.\n\ny = y.astype('str')\ny['id'] += '.tif'\n\nx_train = ImageDataGenerator(\n    validation_split = 0.2\n).flow_from_dataframe(\n    dataframe = y,\n    directory = train_dir,\n    x_col = 'id',\n    y_col = 'label',\n    target_size = (96, 96),\n    class_mode = 'binary',\n    subset = 'training'\n)\n\nx_val = ImageDataGenerator(\n    validation_split = 0.2\n).flow_from_dataframe(\n    dataframe = y,\n    directory = train_dir,\n    x_col = 'id',\n    y_col = 'label',\n    target_size = (96, 96),\n    class_mode = 'binary',\n    subset = 'validation'\n)","metadata":{"execution":{"iopub.status.busy":"2023-04-20T18:52:13.777123Z","iopub.execute_input":"2023-04-20T18:52:13.777966Z","iopub.status.idle":"2023-04-20T19:07:03.958460Z","shell.execute_reply.started":"2023-04-20T18:52:13.777932Z","shell.execute_reply":"2023-04-20T19:07:03.957387Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The following CNN model consists of two convolutional layers with L2 regularization, pooling and dropout layers, and a shallow, batch-normalized dense network. Similar models, including deeper ones, were tested and this architecture was found to have the optimal balance of predictive power and ease of training. Note that in a clinical setting, even greater predictive power would be necessary, and it would be worth exploring much larger architectures. The below model would still represent a reasonable starting point.\n\nThe error produced when training the model with the dropout layer is a known bug in Keras and does not influence model performance.","metadata":{}},{"cell_type":"code","source":"model = Sequential()\n\nmodel.add(Conv2D(\n    filters = 64,\n    kernel_size = 3,\n    activation = 'relu',\n    kernel_regularizer = L2()\n))\nmodel.add(Conv2D(\n    filters = 32,\n    kernel_size = 3,\n    activation = 'relu',\n    kernel_regularizer = L2()\n))\nmodel.add(MaxPooling2D())\nmodel.add(Dropout(0.2))\nmodel.add(Flatten())\n\nmodel.add(Dense(units = 16))\nmodel.add(BatchNormalization())\nmodel.add(Activation('relu'))\nmodel.add(Dense(units = 1, activation = 'sigmoid'))\n\nmodel.compile(\n    optimizer = 'RMSprop',\n    loss = 'binary_crossentropy',\n    metrics = [AUC()]\n)\n\nmodel.fit(\n    x_train,\n    validation_data = x_val,\n    epochs = 3\n)","metadata":{"execution":{"iopub.status.busy":"2023-04-20T19:07:03.959820Z","iopub.execute_input":"2023-04-20T19:07:03.960404Z","iopub.status.idle":"2023-04-20T19:49:48.415622Z","shell.execute_reply.started":"2023-04-20T19:07:03.960363Z","shell.execute_reply":"2023-04-20T19:49:48.414648Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"With a private score AUC of 0.87 (see leaderboard), the model performs reasonably well, considering its small size. In training, it was obvious that including more convolutional layers or more filters improved performance, once overfitting was controlled. The improvements were quite marginal, however.\n\nOne particular element that quite drastically improved performance is the batch-normalization layer. This increased AUC by more than 20%. Applying normalization to the activations of convolutional layers had essentially no effect. This is to be expected, because  *tf.keras.preprocessing.image.ImageDataGenerator()* normalizes the data prior to training.","metadata":{}},{"cell_type":"code","source":"test_dir = top_dir + 'test/'\n\ny_test = pd.read_csv(top_dir + 'sample_submission.csv')\ny_test = y_test.astype('str')\ny_test['id'] += '.tif'\n\n# Ensure shuffle = False\nx_test = ImageDataGenerator(\n).flow_from_dataframe(\n    dataframe = y_test,\n    directory = test_dir,\n    x_col = 'id',\n    y_col = None,\n    target_size = (96, 96),\n    class_mode = None,\n    shuffle = False\n)","metadata":{"execution":{"iopub.status.busy":"2023-04-20T19:49:48.417746Z","iopub.execute_input":"2023-04-20T19:49:48.418498Z","iopub.status.idle":"2023-04-20T19:53:00.793645Z","shell.execute_reply.started":"2023-04-20T19:49:48.418458Z","shell.execute_reply":"2023-04-20T19:53:00.792574Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"pred_labels = model.predict(x_test)","metadata":{"execution":{"iopub.status.busy":"2023-04-20T19:53:00.795184Z","iopub.execute_input":"2023-04-20T19:53:00.796081Z","iopub.status.idle":"2023-04-20T20:01:34.000967Z","shell.execute_reply.started":"2023-04-20T19:53:00.796040Z","shell.execute_reply":"2023-04-20T20:01:33.999868Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"submit = pd.DataFrame()\nsubmit['id'] = y_test['id'].str.partition('.')[0]\nsubmit['label'] = pred_labels[:, 0]\nsubmit.to_csv('submit.csv', index = False)","metadata":{"execution":{"iopub.status.busy":"2023-04-20T20:01:34.002713Z","iopub.execute_input":"2023-04-20T20:01:34.003083Z","iopub.status.idle":"2023-04-20T20:01:34.207966Z","shell.execute_reply.started":"2023-04-20T20:01:34.003035Z","shell.execute_reply":"2023-04-20T20:01:34.206891Z"},"trusted":true},"execution_count":null,"outputs":[]}]}