{"cells":[{"metadata":{"_uuid":"29cef868b4828df297d5d4bfeef48a8ef8d81559"},"cell_type":"markdown","source":"This kernel is my first try at making a NN using keras to apply it to the cancer cell competion. Any comments are more than welcome on any topic, as I am a very early beginner in data science :-)"},{"metadata":{"_uuid":"124276c1d04cc102370f3c303e7f325d9a4b835f"},"cell_type":"markdown","source":"**Things that I learned**\nThis is my first deep learning code, so obviously, it can only be a learning experience. Since I've read a few Kernels already, I figured if any other beginner like me stumbles upon my Kernel, maybe this might be helpful. If not, well at least I get to write feedback for myself ;-)\n\n- I started by trying to make a [57k, 96, 96, 3] np.ndarray containing all the arrays of all the images we need to classify. While that did seem to work with smaller set (I tried with 25k, and it worked), at 57k, the Kernel just crashes. After some investigation (*puts Sherlock's hat on*) the issue seems to be memory overload. I mean I'm just turning 57,000+ images into 96x96x3 arrays, what could go wrong? Next step is to try inserting the prediction inside the for loop. Here's the idea: I'm still training my model (Well, not really mine, rather the one used in the week 2 of Course 4 of Andrew Ng's Coursera Deep Learning course) with a small amount of data, just to see if it's working. I'm taking baby steps, I'll gradually add more data as things work (eventually, I hope). I saw a tweet from Andrej Karapathy a fw weeks ago saying that you should try making a small model, with little data, until it overfits, and then move on. This allows you to check that the model is working, and helps find potential (as a beginner, I'd change the word 'potential' by 'numerous', but maybe it's just me) sources of error. Once I have my first model, then the prediciton is done image by image: I take an image, convert it to an array, and then predict it. Repeat ~57,000 times. This means no storing of the arrays, and (hopefully) no memory crash.\n\n- I did the above, successfully. Meaning that I can get an output and send it for classification (I got 0.6520 with about 2,000 training examples, 4 epochs; and this grew to 0.6793 with 25,000 training examples and 20 epochs). So next step, is to change the model, or neural network. It seems like I am getting only slight improvements with adding more training examples, which is good, but I am sure I can make bigger steps, without using more than a few thousand training examples. Then, once I found a much better model, I will add more examples, and training epochs. Next step: trying the Keras build-in ResNet50."},{"metadata":{"_uuid":"261a9e07c86f65cb749bd65eca0d5e9ae0dfd05b"},"cell_type":"markdown","source":"**TO DO**\nProblem at the moment: memory overload.\n- Train model with about 2000 train ex, 500 test ex, and then do the prediction inside the loop, so no array storing of the image on submission set. In submission loop: img_to_array, then predict, then put prediction into sample_submission['label']. I think this is the file that has to be submitted again, have to check that.\n\n- I have taken the predicitons and made that any prediction >= 0.5 should be considered as 1, and any prediciton < 0.5 should be labelled as 0. I don't know if that 0.5 limit is good, or if it should be higher/lower?\n\n- The competition says that the goal of the NN is to determine if there are cancer pixels in the 32x32 main area of the image. One idea would be to crop the images down to the 32x32 main part, and run the model & estimate, to see if it runs better like that.\n\n- I am using an Adam optimizer, with binary crossentropy loss. This might be interesting to look into aswell"},{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true,"scrolled":true},"cell_type":"code","source":"import numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\nimport cv2\nimport matplotlib.pyplot as plt\nfrom PIL import Image\n\nimport time\n\nimport os\nprint(os.listdir(\"../input\"))","execution_count":null,"outputs":[]},{"metadata":{"_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","trusted":true},"cell_type":"code","source":"train_labels = pd.read_csv('../input/histopathologic-cancer-detection/train_labels.csv')\ntest_labels = pd.read_csv('../input/histopathologic-cancer-detection/sample_submission.csv')\n\n#print('train : ','\\n', train_sample.head(5))\n#print('test : ','\\n', test_labels.head(5))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"1979b9a7c05f5b15aa6d84a4c293b07a431dfc2b"},"cell_type":"code","source":"test_labels.shape","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"87213a83d7f9f262cbae6dd88f15bd62ec45edf2"},"cell_type":"markdown","source":"Very early remarks:\n\n1. We have 220,025 images in train_labels, of which 89,117 are labelled as having a cancer pixel in the 32x32 center zone.\n2. There are 57,5k images in sample_sumbissions\n3. Images are of size 96x96x3 (Meaning RBG, and of total size 27,648)\n4. The only information we have is, if there is a cancer cell of not (in the 32x32 center zone, according to the data description). We do not know what it looks like, which pixel is identified as the one being the cancer, and what makes or doesn't make a cancer cell. This will make EDA relatively fast in my opinion, because there isn't much information we are going to be abel to look at; apart from looking at 1 labelled pictures, and visually trying to find what looks like the patterns compared to 0 labelled data.\n"},{"metadata":{"trusted":true,"_uuid":"b60f8b6a66b399818f814587de9fe39222e6f972"},"cell_type":"code","source":"#This image is labelled as having a cancer cell.\nimage = plt.imread('../input/histopathologic-cancer-detection/train/c18f2d887b7ae4f6742ee445113fa1aef383ed77.tif')\nplt.imshow(image)\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"299b20258395b4eda78581c42b8c166277554ab1"},"cell_type":"code","source":"image.shape","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"8002524e7a2c7408bdd698bf1c1ea835206e4241"},"cell_type":"markdown","source":"Now we are going to try (emphasis on the word 'try') to make a model based on the exciting keras models"},{"metadata":{"trusted":true,"_uuid":"d6710c242ef40e2bc264578f337f15f421f8dfc6"},"cell_type":"code","source":"#let's start with a small sample first:\n\n#size train sample:\nx = 30000\n\n#size of val sample:\nl = 5000\n\ntrain_sample = train_labels[:x]\nval_sample = train_labels[x:x+l]\ntest_sample = test_labels","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"6d6344f596c30427616f56d6de407d51e33a9299"},"cell_type":"code","source":"from keras.preprocessing.image import ImageDataGenerator, array_to_img, img_to_array, load_img\nfrom keras import layers\nfrom keras.layers import Input, Dense, Activation, ZeroPadding2D, BatchNormalization, Flatten, Conv2D\nfrom keras.layers import AveragePooling2D, MaxPooling2D, Dropout, GlobalMaxPooling2D, GlobalAveragePooling2D\nfrom keras.models import Model, Sequential\nfrom keras.optimizers import Adam\nfrom keras.applications.resnet50 import ResNet50\nfrom keras.preprocessing import image\nfrom keras.utils import layer_utils, to_categorical\nfrom keras.utils.data_utils import get_file\nfrom keras.applications.imagenet_utils import preprocess_input\nimport pydot\nfrom IPython.display import SVG\nfrom keras.utils.vis_utils import model_to_dot\nfrom keras.utils import plot_model\nimport tensorflow as tf\nfrom sklearn.metrics import roc_auc_score\n\nimport keras.backend as K\nK.set_image_data_format('channels_last')\nimport matplotlib.pyplot as plt\nfrom matplotlib.pyplot import imshow\n\n%matplotlib inline","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"df218a9957e828be58f23003240bf728faf96d61"},"cell_type":"code","source":"img_heigth, img_width = 96, 96","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"37275d48806261b1249c34e4d4625e93c2938054"},"cell_type":"code","source":"img=load_img('../input/histopathologic-cancer-detection/train/c18f2d887b7ae4f6742ee445113fa1aef383ed77.tif')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"6ed650863f8ecc96cb9642c7afe2375cdd95ac24"},"cell_type":"code","source":"train_sample.iloc[0][1]","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"91f666c49965788657804c390d7d8a1445701693"},"cell_type":"code","source":"nb_train_examples=train_sample.shape[0]\nnb_val_examples=val_sample.shape[0]\n\ntrain_img_array = np.ndarray(shape=[nb_train_examples, 96, 96, 3])\ntrain_img_label = np.ndarray(shape=[nb_train_examples, 1])\n\nval_img_array = np.ndarray(shape=[nb_val_examples, 96, 96, 3])\nval_img_label = np.ndarray(shape=[nb_val_examples, 1])\n\ntest_img_array = np.ndarray(shape=[test_sample.shape[0], 96, 96, 3])\ntest_img_label = np.ndarray(shape=[test_sample.shape[0], 1])","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"705d19729b5acc3139fdd6665d6b544c145b1248"},"cell_type":"code","source":"t1=time.time()\nfor p in range(nb_train_examples):\n    #We turn the .tif into an array\n    img_name=train_sample.iloc[p][0]\n    img=load_img('../input/histopathologic-cancer-detection/train/'+img_name+'.tif')\n    img=img_to_array(img)\n    img=img/255\n    #print(img_name)\n    #print(img.shape)\n    train_img_array[p]=img #putting the image inside the 4 dim array\n    \n    #We put the label into a new ndarray:\n    train_img_label[p]=train_sample.iloc[p][1]\nt2=time.time()\nprint('time to turn .tif into array for train_set : ',t2-t1)\nprint('train_img_array shape is : ', train_img_array.shape)\nprint('train_img_label shape is : ', train_img_label.shape)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"fc1f7a77f668378c40ed6a633f7fca989dab46f1"},"cell_type":"code","source":"t1=time.time()\nfor p in range(nb_val_examples):\n    #We turn the .tif into an array\n    img_name=val_sample.iloc[p][0]\n    img=load_img('../input/histopathologic-cancer-detection/train/'+img_name+'.tif')\n    img=img_to_array(img)\n    img=img/255\n    #print(img_name)\n    #print(img.shape)\n    val_img_array[p]=img #putting the image inside the 4 dim array\n    \n    #We put the label into a new ndarray:\n    val_img_label[p]=val_sample.iloc[p][1]\nt2=time.time()\nprint('time to turn .tif into array for val_sample : ',t2-t1)\nprint('val_img_array shape is : ', val_img_array.shape)\nprint('val_img_label shape is : ', val_img_label.shape)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"073756bf6dd325770c607da1673d13718e206f1e"},"cell_type":"markdown","source":"It takes about 2.42 seconds to do img_to_array on 1,000 examples. This means it's going to take about 8.4 min for all 220k training examples, + 2,3 min for the test set. That's about 11 min in total just to convert the images to numpy using this method. (goes much much faster whe using GPU)\n\nIt's long, but at least it works, so for now, that's good enough."},{"metadata":{"trusted":true,"_uuid":"1d49c6561c9904d0c404ae955f89995defac413b"},"cell_type":"code","source":"def auroc(y_true, y_pred):\n    return tf.py_func(roc_auc_score, (y_true, y_pred), tf.double)\n\ndef FirstModel(num_classes):\n    \n    model = Sequential()\n    model.add(ResNet50(include_top = False, pooling='avg'))\n    model.add(Dense(num_classes, activation = 'sigmoid'))\n    \n    model.layers[0].trainable = False\n    \n    return model","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"2fda2268498c9b8331e4b6f9d7d35e49c2838388","scrolled":false},"cell_type":"code","source":"#my_model=FirstModel(train_img_array[0].shape)\nmy_model = FirstModel(num_classes = 2)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"a145bad784beb3c40496f69a4582e94da504e450"},"cell_type":"code","source":"my_model.compile(optimizer = Adam(lr=0.0001), loss = 'binary_crossentropy', metrics = ['accuracy', auroc])","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"2c833fa18cb03029918247b00e928e6129c052af"},"cell_type":"code","source":"train_img_label = to_categorical(train_img_label, num_classes=2)\nval_img_label = to_categorical(val_img_label, num_classes=2)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"scrolled":true,"_uuid":"02ab0857b14ee0456143d4678174b1b43b8ea4e3"},"cell_type":"code","source":"stats = my_model.fit(x = train_img_array, y = train_img_label, epochs = 5)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"a29a06c87cbbfbb0e6a190e4ec86335e667e8ab8"},"cell_type":"code","source":"evaluation = my_model.evaluate(x= val_img_array, y=val_img_label)\nprint()\nprint (\"Loss = \" + str(evaluation[0]))\nprint (\"Test Accuracy = \" + str(evaluation[1]))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"92de0b1488dd90e2fccf117ab3c33885726d34bb"},"cell_type":"code","source":"#This turned out to be a bad idea. But I'm keeping it, never know when I might need it\n\ndummy_img = np.ndarray(shape=(1, 96, 96, 3))\n\nt1=time.time()\nfor p in range(test_sample.shape[0]):\n    #We turn the .tif into an array\n    img_name=test_sample.iloc[p][0]\n    img=load_img('../input/histopathologic-cancer-detection/test/'+img_name+'.tif')\n    img=img_to_array(img)\n    img=img/255\n    \n    #print(img_name)\n    #print(img.shape)\n    \n    pred = my_model.predict(img.reshape(1,96,96,3))\n    #We put the label into a new ndarray:\n    test_sample.at[p ,'label'] = pred.argmax()\n    \n    \nt2=time.time()\nprint('time to turn .tif into array for test_sample : ',t2-t1)\n","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"675073fba19507d889965c9e717cecf2b4c8273d"},"cell_type":"code","source":"print('number of images labelled with cancer : ',test_sample[test_sample['label']==1].shape[0],\n      ' out of ', test_sample.shape[0], ' examples')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"df690179fc56b9f02978ad811d094e8ef28a0744"},"cell_type":"code","source":"test_sample.to_csv('test_predictions.csv', index=False)","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.6.6","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"}},"nbformat":4,"nbformat_minor":1}