{"cells":[{"metadata":{"_uuid":"f3663513761715e56d3f6f27e3e0911c99e4432d"},"cell_type":"markdown","source":"# Introduction\n____\nIn this kernel I'll show how to get the data in the proper format to use it on a CNN and make the classification. First let's import the necessary libraries."},{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true},"cell_type":"code","source":"import numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\nfrom glob import glob \nfrom skimage.io import imread #read images from files\nimport os\nimport keras.backend as K\nimport tensorflow as tf","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"c015810a1def64d5b604f9625f986d8abe56c562"},"cell_type":"markdown","source":"Next we create a dataframe using the image's path as the first column and the id as the second column (if you don't understand why that split is being used that way, I suggest that you get one of the paths available and try using `.split('/')[3]...` to see what is going on here. Next we read the labels and merge with our dataframe through their ids, so we know which image corresponds to each label. "},{"metadata":{"_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","trusted":true,"scrolled":false},"cell_type":"code","source":"base_tile_dir = '../input/train/'\ndf = pd.DataFrame({'path': glob(os.path.join(base_tile_dir,'*.tif'))})\ndf['id'] = df.path.map(lambda x: x.split('/')[3].split(\".\")[0])\nlabels = pd.read_csv(\"../input/train_labels.csv\")\ndf = df.merge(labels, on = \"id\")\ndf.head(10)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"e20aa314e20a260650a6aaef935b3398fa8f5696"},"cell_type":"markdown","source":"Now, before we look at the images, it is important to note that there are LOTS of images and this can easily blow up the kernel's memory if we use all of them in training (considering that we are using a Kaggle's kernel). To avoid this problem and also keep balance in our training data, I'll use 5k examples for each label. "},{"metadata":{"trusted":true,"scrolled":true,"_uuid":"ea4099a7bd02db79a467fceb461891df09f5d9c5"},"cell_type":"code","source":"df0 = df[df.label == 0].sample(5000, random_state = 42)\ndf1 = df[df.label == 1].sample(5000, random_state = 42)\ndf = pd.concat([df0, df1], ignore_index=True).reset_index()\ndf = df[[\"path\", \"id\", \"label\"]]\ndf.sample(10)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"588618bc09a0461057dbebd92ab7f48ba6eae730"},"cell_type":"markdown","source":"Now that we know that both classes are balanced and how many of them there are, we can start looking at the images. To do so, I'll use imread function imported on the first code chunk."},{"metadata":{"trusted":true,"_uuid":"6d50c12e2c782c14ff07a14e06e105b96b9f948c"},"cell_type":"code","source":"df['image'] = df['path'].map(imread)\ndf.sample(3)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"b7badd35f32db6c5d38b36e45f9989baea81b746"},"cell_type":"markdown","source":"Great! Now we can see that our images are represented by an array of arrays. To read them, we can make use of either matplotlib or the more computer vision-focused opencv. Here I'll use the first option for simplicity. Let's see two sets of images: one with label 0 and another with label 1."},{"metadata":{"trusted":true,"_uuid":"b607b80f99c86e34b2912ef3730bfe7686b33dd2"},"cell_type":"code","source":"import matplotlib.pyplot as plt\n\nimages = [(df['image'][0], df['label'][0]), \n          (df['image'][1], df['label'][1]),\n          (df['image'][2], df['label'][2]),\n          (df['image'][5000], df['label'][5000]),\n          (df['image'][5001], df['label'][5001]),\n          (df['image'][5002], df['label'][5002])]\n\nfig, m_axs = plt.subplots(1, len(images), figsize = (20, 2))\n#show the images and label them\nfor ii, c_ax in enumerate(m_axs):\n    c_ax.imshow(images[ii][0])\n    c_ax.set_title(images[ii][1])","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"6aae464bd87e7aec2539cf630c22dc1b7243c005"},"cell_type":"markdown","source":"Only by looking at the images above I can hardly tell why the three to the right were detected with cancer. Maybe our model can find patterns that we - that don't have any specialized training - can't? Let's create our set of inputs. By using `np.stack` we create a 4-rank tensor: 10000 observations where each one is a 96x96x3 image."},{"metadata":{"trusted":true,"_uuid":"c0f45b8e8937c05d9acd3ae05cbdae73eab3f346"},"cell_type":"code","source":"input_images = np.stack(list(df.image), axis = 0)\ninput_images.shape","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"9269aad9f331cab20f228a0009a517dc52bfd56f"},"cell_type":"markdown","source":"Now our work is pretty straightforward. We split our input_images in training, validation and testing sets and run them through our model."},{"metadata":{"trusted":true,"scrolled":false,"_uuid":"3efe31cccadd788f56e7a789d31ee05f6ef0ac95"},"cell_type":"code","source":"from sklearn.model_selection import train_test_split\nfrom sklearn.preprocessing import LabelBinarizer\nfrom keras.utils import np_utils\n\n\ntrain_fraction = 0.8\n\nencoder = LabelBinarizer()\ny = encoder.fit_transform(df.label)\nx = input_images\n\ntrain_tensors, test_tensors, train_targets, test_targets =\\\n    train_test_split(x, y, train_size = train_fraction, random_state = 42)\n\nval_size = int(0.5*len(test_tensors))\n\nval_tensors = test_tensors[:val_size]\nval_targets = test_targets[:val_size]\ntest_tensors = test_tensors[val_size:]\ntest_targets = test_targets[val_size:]","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"2c967f26ff7d0bcd68ff2bba9631aa333fcaea1f"},"cell_type":"code","source":"from keras.layers import Conv2D, MaxPooling2D, GlobalAveragePooling2D, GlobalMaxPooling2D\nfrom keras.layers import Dropout, Flatten, Dense\nfrom keras.callbacks import EarlyStopping, ModelCheckpoint\nfrom keras.models import Sequential\nfrom tensorflow import set_random_seed\n\nset_random_seed(42)\n\nearly_stopping = EarlyStopping(monitor = 'val_loss', patience = 5)\ncheckpointer = ModelCheckpoint(filepath='weights.hdf5', \n                               verbose=1, save_best_only=True)\nmodel = Sequential()\nmodel.add(Conv2D(filters = 16, kernel_size = 3, padding = 'same', activation = 'relu', input_shape = (96, 96, 3)))\nmodel.add(Conv2D(filters = 16, kernel_size = 3, padding = 'same', activation = 'relu'))\nmodel.add(Conv2D(filters = 16, kernel_size = 3, padding = 'same', activation = 'relu'))\nmodel.add(Dropout(0.3))\nmodel.add(MaxPooling2D(pool_size = 3)) \n\nmodel.add(Conv2D(filters = 32, kernel_size = 3, padding = 'same', activation = 'relu')) \nmodel.add(Conv2D(filters = 32, kernel_size = 3, padding = 'same', activation = 'relu')) \nmodel.add(Conv2D(filters = 32, kernel_size = 3, padding = 'same', activation = 'relu'))\nmodel.add(Dropout(0.3))\nmodel.add(MaxPooling2D(pool_size = 3)) \n\nmodel.add(Conv2D(filters = 64, kernel_size = 3, padding = 'same', activation = 'relu'))\nmodel.add(Conv2D(filters = 64, kernel_size = 3, padding = 'same', activation = 'relu'))\nmodel.add(Conv2D(filters = 64, kernel_size = 3, padding = 'same', activation = 'relu'))\nmodel.add(Dropout(0.3))\nmodel.add(MaxPooling2D(pool_size = 3))\n\nmodel.add(Conv2D(filters = 128, kernel_size = 3, padding = 'same', activation = 'elu'))\nmodel.add(Conv2D(filters = 128, kernel_size = 3, padding = 'same', activation = 'elu'))\nmodel.add(Conv2D(filters = 256, kernel_size = 3, padding = 'same', activation = 'elu'))\n\nmodel.add(Flatten())\nmodel.add(Dense(1, activation = 'sigmoid'))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"05efe7564db8f82544085978fa319fa1517f8a72","scrolled":false},"cell_type":"code","source":"#working AUC metric for keras from here: \n#https://www.kaggle.com/c/porto-seguro-safe-driver-prediction/discussion/41015\ndef auc(y_true, y_pred):   \n    ptas = tf.stack([binary_PTA(y_true,y_pred,k) for k in np.linspace(0, 1, 1000)],axis=0)\n    pfas = tf.stack([binary_PFA(y_true,y_pred,k) for k in np.linspace(0, 1, 1000)],axis=0)\n    pfas = tf.concat([tf.ones((1,)) ,pfas],axis=0)\n    binSizes = -(pfas[1:]-pfas[:-1])\n    s = ptas*binSizes\n    return K.sum(s, axis=0)\n\n#---------------------\n# PFA, prob false alert for binary classifier\ndef binary_PFA(y_true, y_pred, threshold=K.variable(value=0.5)):\n    y_pred = K.cast(y_pred >= threshold, 'float32')\n    # N = total number of negative labels\n    N = K.sum(1 - y_true)\n    # FP = total number of false alerts, alerts from the negative class labels\n    FP = K.sum(y_pred - y_pred * y_true)\n    return FP/N\n\n#----------------\n# PTA prob true alerts for binary classifier\ndef binary_PTA(y_true, y_pred, threshold=K.variable(value=0.5)):\n    y_pred = K.cast(y_pred >= threshold, 'float32')\n    # P = total number of positive labels\n    P = K.sum(y_true)\n    # TP = total number of correct alerts, alerts from the positive class labels\n    TP = K.sum(y_pred * y_true)\n    return TP/P\n\nfrom keras.optimizers import SGD\n# sgd = SGD(lr = 0.0001, momentum = 0.9)\nmodel.compile(optimizer= 'adam', loss='binary_crossentropy', metrics=['accuracy'])\n\nepochs = 15\nmodel.fit(train_tensors, train_targets, \n          validation_data=(val_tensors, val_targets),\n          epochs=epochs, batch_size=80, verbose=1, callbacks = [early_stopping, checkpointer])","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"60974b631ffd85ba194976f124a3928ee409948d"},"cell_type":"code","source":"model.load_weights('weights.hdf5')\n\ncancer_predictions =  [model.predict(np.expand_dims(tensor, axis=0))[0][0] for tensor in test_tensors]\n\ntest_accuracy = 100*np.sum(np.round(cancer_predictions).astype('int32')==test_targets.flatten())/len(cancer_predictions)\nprint('Test accuracy: %.4f%%' % test_accuracy)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"f2ddaeb2d573d33d9270244387adb2d0fc6c5929"},"cell_type":"code","source":"#AUC score\nfrom sklearn.metrics import roc_auc_score\nscore = roc_auc_score(np.round(cancer_predictions).astype('int32'), test_targets)\nscore","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"76097be9c4920e847c1595ac967f4e79080dbb20"},"cell_type":"markdown","source":"Well, the nest accuracy with this kernel is around 80% which is pretty good for such a straightfoward model. And here is the magic of Deep Learning. Without any feature preprocessing nor previous knowledge, we were able to correctly classify around 800 out of 1000 images. Now it's time to make our predictions and submit the results. \n\nNote: The accuracies may be lower than 80% due to some randomness. I'm using previously set random seeds where I can, but I've read that there can be some randomness when running in GPU (which I am for faster kernels) and this makes things harder to follow. However I think this is a good starting point for anyone that wants to get into CNN using keras. You can later check other kernels in this competition that use transfer learning from previously trained neural nets like NasNet, Resnet, Inception, Xception etc.\n\nThe preprocessing steps are very similar to what we did before in the training set."},{"metadata":{"trusted":true,"_uuid":"e3c166fb81a5741b192074b92c6373c7461db9e4"},"cell_type":"code","source":"base_tile_dir = '../input/test/'\ntest_df = pd.DataFrame({'path': glob(os.path.join(base_tile_dir,'*.tif'))})\ntest_df['id'] = test_df.path.map(lambda x: x.split('/')[3].split(\".\")[0])","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"4770bab75f46837ef793da11fef8a51bb66046f4"},"cell_type":"code","source":"test_df['image'] = test_df['path'].map(imread)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"scrolled":false,"_uuid":"caa208d0bf9fe841fd105bf65e0948e26110b5b3"},"cell_type":"code","source":"test_images = np.stack(test_df.image, axis = 0)\ntest_images.shape","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"d4cfab07135ac0cf5237297604e457ad1452a94a"},"cell_type":"markdown","source":"Using the trained model to make the predictions"},{"metadata":{"trusted":true,"_uuid":"efad9d1287c1c61cc52f959814ced4fe696be9d5"},"cell_type":"code","source":"predicted_labels =  [model.predict(np.expand_dims(tensor, axis=0))[0][0] for tensor in test_images]","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"32832b729dfefaeb2368fa0481698ebb5d8b3f66"},"cell_type":"code","source":"predictions = np.array(predicted_labels)\ntest_df['label'] = predictions\nsubmission = test_df[[\"id\", \"label\"]]\nsubmission.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"dcfd06342c511340de1980df1243a082da394f2c"},"cell_type":"code","source":"#submission\nsubmission.to_csv(\"submission.csv\", index = False, header = True)","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.6.6","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"}},"nbformat":4,"nbformat_minor":1}