{"cells":[{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true},"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load in \n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n\n# Input data files are available in the \"../input/\" directory.\n# For example, running this (by clicking run or pressing Shift+Enter) will list the files in the input directory\n\nimport os\nprint(os.listdir(\"../input\"))\n\n# Any results you write to the current directory are saved as output.","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"35630c99d23092ae5f150c0efb6b8d450f53d405"},"cell_type":"markdown","source":"*** 1. READING IN THE FILES**"},{"metadata":{"_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","trusted":true},"cell_type":"code","source":"#read the training data file\ndf = pd.read_csv('../input/train.csv')\nprint(df.head())","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"abcc6aae336f13800a95c8585479c4e6ed52b3f8"},"cell_type":"markdown","source":"> Here we have two columns. One colum corresponds to the images for the whales while the other column corresponds to the image id for the whale."},{"metadata":{"trusted":true,"_uuid":"932adca757754444ee20022ca059f9556eda714b"},"cell_type":"code","source":"#let's look at the unique values in the ID column\nprint(df['Id'].describe())","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"3fe8b770be31edea3f5af34093fe912f61ffdd58"},"cell_type":"code","source":"import seaborn as sns\nimport matplotlib.pyplot as plt\nplt.figure(figsize = (15,10))\nsns.countplot(y = df['Id'] == 'new_whale', palette = 'Dark2')","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"7394f1755a7c80d9268b3b9a76e7c34b56c51789"},"cell_type":"markdown","source":">From the above plot we can see that the count of ID's corresponding to the pics that were labelled as new_whale is maximum. The images with new whales amount to a total of about 9700 which is more than half of total available images."},{"metadata":{"_uuid":"8eac664c0d063577e5b5ed51d1548f15d9a497e2"},"cell_type":"markdown","source":"> **Let's prepare our data so that we can train our CNN model onto it. The images here are in the form of string. We know that our CNN model takes images in the form of array as input. So we need to convert our string into array format so that we can feed them to our CNN.**"},{"metadata":{"trusted":true,"_uuid":"d90e191da384ff78c16787ad05fd3b6f9c957ec4"},"cell_type":"code","source":"#dimension of our original training dataframe\nprint(df.shape)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"8d4624cdc5c47e4d9cbd935d4c13029d498a788c"},"cell_type":"code","source":"#Let's get our x_train and y_train from our dataframe\nx_train = df['Image']\ny_train = df['Id']","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"d0413157e48dc4a63d71c74e9e6fc7c6befb4734"},"cell_type":"markdown","source":"**2. PREPARING OUR TRAINING IMAGE DATA**"},{"metadata":{"_uuid":"c147f31b5b465d9dddf4f71d80ecb7e03bff2329"},"cell_type":"markdown","source":">**while obtaining our data into an array format our data has to follow a specific format. The usual format for the input data that is to be feed to the CNN model is of the form --> (batch_size, height, width, channels). Here the batch size is nothing but the total number of rows in our train dataframe i.e 25361. The height and width is something that we can arbitrarily choose. It's always a good practice to choose the height and width in such a way as it should be as minimum as possible but not as small that it would become difficult for us to interpret the image itself. So the values should be such that we should be able to tell the content of the image by looking at it. The main idea is to reduce the total number of parameters as far as possible to reduce the computational time.**"},{"metadata":{"trusted":true,"_uuid":"0017cdc63c4839714d1f0da5c41b107da78dc95f"},"cell_type":"code","source":"#import all the necessary libraries from the keras API\nimport keras\nfrom keras.preprocessing.image import load_img\nfrom keras.preprocessing.image import img_to_array\nfrom keras.applications.imagenet_utils import preprocess_input","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"2c3cb841cff23eea39323831ad8bf723e02e70fc"},"cell_type":"code","source":"#define a function to prepare our trianing images\ndef PrepareTrainImages(dataframe, shape, path):\n    \n    #obtain the numpy array filled with zeros having the format --> (batch_size, height, width, channels)\n    x_train = np.zeros((shape, 100, 100, 3))\n    count = 0\n    \n    for fig in dataframe['Image']:\n        \n        #load images into images of size 100x100x3\n        img = load_img(\"../input/\" + path + \"/\" + fig, target_size = (100, 100, 3))\n        \n        #convert images to array\n        x = img_to_array(img)\n        x = preprocess_input(x)\n\n        x_train[count] = x\n        count += 1\n    \n    return x_train\n    ","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"31f87dc9555ed8a5b1a9f4f0864c72aa1f7ce911"},"cell_type":"code","source":"x_train = PrepareTrainImages(df, df.shape[0], 'train')","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"47676afd2c3cb0b0f2f3d09ae838bbb4ef48f28f"},"cell_type":"markdown","source":"**3. NORMALIZE THE DATA**"},{"metadata":{"_uuid":"b6f17ea7ab909f594f0c5231357db1c24d636430"},"cell_type":"markdown","source":"**Once we got the training data the next step would be to normalize the data so that all the pixel values lie in the same range.**"},{"metadata":{"trusted":true,"_uuid":"fb453972f1a969c04d4bcbf75e1e2a17d8e83d84"},"cell_type":"code","source":"print(x_train.shape) #we got the data in the format that we need for the CNN model","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"3a8e62b0262dd06e633aa96b2c49540ebae44a99"},"cell_type":"code","source":"#let's normalize the data.\nx_train[0] # we can see that the pixel values in the following array have large differene in their values\n#so it's always better the obtain all the values in the same range\nx_train = x_train.astype('float32') / 255 #data normalized","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"ebcfbab5266cdb8d4da5775f87d39104a48dc194"},"cell_type":"markdown","source":"**4. DATA VISUALIZATION**"},{"metadata":{"trusted":true,"_uuid":"4b057f2ef248b0a0b78bb0d04ca07b94e8544ea8"},"cell_type":"code","source":"#Let's visualize some of our taining images\nplt.figure(figsize = (12,8))\nplt.subplot(2, 2, 1)\nplt.imshow(x_train[0][:,:,0], cmap = 'gray') #the first image\nplt.title(df.iloc[0,0])\nplt.xticks([])\nplt.yticks([])\n\nplt.subplot(2, 2, 2)\nplt.imshow(x_train[100][:,:,0], cmap = 'gray')\nplt.title(df.iloc[100,0])\nplt.xticks([])\nplt.yticks([])\n\nplt.subplot(2, 2, 3)\nplt.imshow(x_train[1000][:,:,0], cmap = 'gray')\nplt.title(df.iloc[1000,0])\nplt.xticks([])\nplt.yticks([])\n\nplt.subplot(2, 2, 4)\nplt.imshow(x_train[4000][:,:,0], cmap = 'gray')\nplt.title(df.iloc[4000,0])\nplt.xticks([])\nplt.yticks([])","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"07b28bde5af75e51d5b2fcdbf0286ae47ff9d796"},"cell_type":"markdown","source":"**5. LABLE ENCODING AND ONE HOT ENCODING THE ID COLUMN  VALUES**"},{"metadata":{"trusted":true,"_uuid":"b1661c0365f231307b28dd671fa4e8b4e601dc2a"},"cell_type":"code","source":"from keras.utils import np_utils #to obtain the one hot encodings of the id values\nfrom sklearn.preprocessing import LabelEncoder #to obtain the unique integer values for each id values","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"89b26ba9f69865133c0ac7b73d55d0b95d7fe0e1"},"cell_type":"code","source":"le = LabelEncoder()\ny_train = np_utils.to_categorical(le.fit_transform(y_train))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"60e61ba6f0d49de448ef0b657661e7c164455ae4"},"cell_type":"code","source":"print(y_train[:10])\nprint(y_train.shape)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"d97166b35606fd5f5bab00ad45e281fc51ad0b16"},"cell_type":"markdown","source":"**6. BUILDING THE CNN MODEL**"},{"metadata":{"_uuid":"e909130d76f292e522d30cea3dc6222fae75ae1e"},"cell_type":"markdown","source":"**Now that we have our training dataset prepared we are ready to build our CNN model. Let's start by building a simple CNN model\nThis model will have the following layer arrangements**\n\n***(CONV2D -> ACTIVATION -> MAXPOOLING) --- (CONV2D -> ACTIVATION -> MAXPOOLING) --- (FLATTEN)---(DENSE -> ACTIVATION)***"},{"metadata":{"trusted":true,"_uuid":"9fa1d6397e0e3bce76d8571900a25bed3e3c1e51"},"cell_type":"code","source":"#let's start by importing all the necessary libraries for building the CNN model\nimport keras\nfrom keras.layers import Conv2D\nfrom keras.layers import Activation, BatchNormalization\nfrom keras.layers import MaxPooling2D, Dropout\nfrom keras.layers import Flatten, Dense\nfrom keras.models import Sequential\nfrom keras.optimizers import Adam","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"68f70abe7fa0f99af4a20b1f57d5025faa0cd5f4"},"cell_type":"markdown","source":">**We will start by building the CNN model without implementing any regularization techniques such as dropout or batch normalization just to analyze how important it is to use some of the regularization techniques almost everytime to avoid overfitting of your model."},{"metadata":{"trusted":true,"_uuid":"f4dbad4af45e6677f2def40286378bd275998065"},"cell_type":"code","source":"#start building the model\nmodel = Sequential()\nmodel.add(Conv2D(32, (3,3), input_shape = (x_train.shape[1:]), padding = 'same'))\nmodel.add(Activation('relu'))\nmodel.add(MaxPooling2D(pool_size =  (2,2)))\n\nmodel.add(Conv2D(64, (3,3), padding = 'same'))\nmodel.add(Activation('relu'))\nmodel.add(MaxPooling2D(pool_size =  (2,2)))\n\nmodel.add(Conv2D(128, (3,3), padding = 'same'))\nmodel.add(Activation('relu'))\nmodel.add(MaxPooling2D(pool_size =  (2,2)))\n\nmodel.add(Flatten())\n\nmodel.add(Dense(512))\nmodel.add(Activation('relu'))\n\nmodel.add(Dense(y_train.shape[1]))\nmodel.add(Activation('softmax'))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"481d2d08993256ddd09f696d233e747da36776d6"},"cell_type":"code","source":"#looking at the summary for our model\nmodel.summary()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"b27f56f99c9b559017e890320fc219746122ca4d"},"cell_type":"code","source":"#compile the model\noptim = Adam(lr = 0.001) #using the already available learning rate scheduler\nmodel.compile(loss = 'categorical_crossentropy', optimizer = optim, metrics = ['accuracy'])","execution_count":null,"outputs":[]},{"metadata":{"_kg_hide-input":true,"_kg_hide-output":true,"trusted":true,"_uuid":"686af21358cdd4b4712188a5e3d5a12a78ed94d8"},"cell_type":"code","source":"#fit the model on our dataset\nhistory = model.fit(x_train, y_train, epochs = 30, batch_size = 64)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"7696ec4ad4384ecdef6e8b55f6f05a75cae0a5d1"},"cell_type":"code","source":"#let's look how our model performed by plotting the accuracy and loss curves\nsns.set(style = 'darkgrid')\nplt.figure(figsize = (18, 14))\nplt.subplot(2, 1, 1)\nplt.plot(range(30), history.history['acc'])\nplt.xlabel('EPOCHS')\nplt.ylabel('TRAINING ACCURACY')\nplt.title('TRAINING ACCURACY vs EPOCHS')\n\nplt.figure(figsize = (18, 14))\nplt.subplot(2, 1, 1)\nplt.plot(range(30), history.history['loss'])\nplt.xlabel('EPOCHS')\nplt.ylabel('TRAINING LOSS')\nplt.title('TRAINING LOSS vs EPOCHS')","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"cc393fbffd5c5589c134153e4378cac360c5e4a7"},"cell_type":"markdown","source":"**CNN WITH BATCH NORMALIZATION**"},{"metadata":{"_uuid":"95e65c25a201d3797c45a793fa58c94df9181c02"},"cell_type":"markdown","source":">Here I haven't used maxpooling for downsampling. Instead I have increased the amount of stride each (5,5) filter will take while moving over the input image. Also Batch normalization is implemented as a regularization technique to avoid overfitting."},{"metadata":{"trusted":true,"_uuid":"612a08efa85591e6f0e7d4b14bd00773b3894efc"},"cell_type":"code","source":"#start building the model\nmodel1 = Sequential()\nmodel1.add(Conv2D(32, (5,5), strides = (2,2), input_shape = (x_train.shape[1:]), padding = 'same'))\nmodel1.add(Activation('relu'))\nmodel1.add(BatchNormalization())\n#model1.add(MaxPooling2D(pool_size =  (2,2)))\n\nmodel1.add(Conv2D(32, (3,3), strides = (2,2), padding = 'same'))\nmodel1.add(Activation('relu'))\nmodel1.add(BatchNormalization())\n#model1.add(MaxPooling2D(pool_size =  (2,2)))\n\n# model1.add(Conv2D(128, (3,3), padding = 'same'))\n# model1.add(Activation('relu'))\n# model1.add(BatchNormalization())\n# model1.add(MaxPooling2D(pool_size =  (2,2)))\n\nmodel1.add(Flatten())\n\nmodel1.add(Dense(32))\nmodel1.add(Activation('relu'))\nmodel1.add(BatchNormalization())\n\nmodel1.add(Dense(y_train.shape[1]))\nmodel1.add(Activation('softmax'))\n\nmodel1.summary()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"f6d2eed7623c53cb6c0dc6563fee1862c3571a89"},"cell_type":"code","source":"#compile the model\noptim = Adam(lr = 0.001) #using the already available learning rate scheduler\nmodel1.compile(loss = 'categorical_crossentropy', optimizer = optim, metrics = ['accuracy'])","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"98fee92ba2f17fc7e1307e06a4c45304f276e024"},"cell_type":"code","source":"#fit the model on our dataset\nhistory1 = model1.fit(x_train, y_train, epochs = 25, batch_size = 64)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"5a93c4a91363439cc4c510999c016e91a9c548fd"},"cell_type":"code","source":"#let's look how our model performed by plotting the accuracy and loss curves\nsns.set(style = 'darkgrid')\nplt.figure(figsize = (18, 14))\nplt.subplot(2, 1, 1)\nplt.plot(range(25), history1.history['acc'])\nplt.xlabel('EPOCHS')\nplt.ylabel('TRAINING ACCURACY')\nplt.title('TRAINING ACCURACY vs EPOCHS')\n\nplt.figure(figsize = (18, 14))\nplt.subplot(2, 1, 1)\nplt.plot(range(25), history1.history['loss'])\nplt.xlabel('EPOCHS')\nplt.ylabel('TRAINING LOSS')\nplt.title('TRAINING LOSS vs EPOCHS')","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"7792ae2652b45b7c6ae1010b9e4558ea72104471"},"cell_type":"markdown","source":"**CNN WITH BATCH NORMALIZATION AND DROPOUT**"},{"metadata":{"_uuid":"ad47994761966256ba7abee9e853753823982e96"},"cell_type":"markdown","source":"> Here instead of increasing the strides I have used Maxpooling for downsampling. Also dropout and Batch normalization is also been used as regularization techniques."},{"metadata":{"trusted":true,"_uuid":"fa1648d05eab5a6b0f479f4ece3bc3b5d0245aa5"},"cell_type":"code","source":"#start building the model\nmodel2 = Sequential()\nmodel2.add(Conv2D(32, (5,5), input_shape = (x_train.shape[1:]), padding = 'same'))\nmodel2.add(Activation('relu'))\nmodel2.add(BatchNormalization())\nmodel2.add(MaxPooling2D(pool_size =  (2,2), strides = (2,2)))\nmodel2.add(Dropout(0.2))\n\nmodel2.add(Conv2D(32, (3,3), padding = 'same'))\nmodel2.add(Activation('relu'))\nmodel2.add(BatchNormalization())\nmodel2.add(MaxPooling2D(pool_size =  (2,2), strides = (2,2)))\nmodel2.add(Dropout(0.2))\n\n# model1.add(Conv2D(128, (3,3), padding = 'same'))\n# model1.add(Activation('relu'))\n# model1.add(BatchNormalization())\n# model1.add(MaxPooling2D(pool_size =  (2,2)))\n\nmodel2.add(Flatten())\n\nmodel2.add(Dense(128))\nmodel2.add(Activation('relu'))\nmodel2.add(BatchNormalization())\nmodel2.add(Dropout(0.5))\n\nmodel2.add(Dense(y_train.shape[1]))\nmodel2.add(Activation('softmax'))\n\nmodel2.summary()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"34d075c56593202f4beb804bf6d85fdb2be1f32d"},"cell_type":"code","source":"#compile the model\noptim = Adam(lr = 0.001) #using the already available learning rate scheduler\nmodel2.compile(loss = 'categorical_crossentropy', optimizer = optim, metrics = ['accuracy'])","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"594bf59f4a3ce3d9bf242e4d2c91b2ba4735703d"},"cell_type":"code","source":"#fit the model on our dataset\nhistory2 = model2.fit(x_train, y_train, epochs = 100, batch_size = 64)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"037d193b2efbf0d4ce466f31250ac47281499c1a"},"cell_type":"code","source":"#let's look how our model performed by plotting the accuracy and loss curves\nsns.set(style = 'darkgrid')\nplt.figure(figsize = (18, 14))\nplt.subplot(2, 1, 1)\nplt.plot(range(100), history2.history['acc'])\nplt.xlabel('EPOCHS')\nplt.ylabel('TRAINING ACCURACY')\nplt.title('TRAINING ACCURACY vs EPOCHS')\n\nplt.figure(figsize = (18, 14))\nplt.subplot(2, 1, 1)\nplt.plot(range(100), history2.history['loss'])\nplt.xlabel('EPOCHS')\nplt.ylabel('TRAINING LOSS')\nplt.title('TRAINING LOSS vs EPOCHS')","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"1ae0839bbb0982c404219557b7b9be4c4cb4f5d5"},"cell_type":"markdown","source":"> If we try to analyze the above three different CNN models, we can clearly see some difference between the model using a regularization technique and the one which isn't. The first model in which no regularization technique was implemented resulted in a good accuracy but to reach that accuracy we required very little time. This is not always bad but also not ideal. The model must take it's time to learn all the features from our dataset and as the epochs progresses the model starts to learn more and more about the dataset and the loss starts to gradually decrease. In the first model the loss did decrease but the decrease was very steep. While in the third model where we implemented both dropout as well as batch normalization the decrease in loss was gradual which is what it should be. While for the model one reached good accuracy in less number of epochs it took more number of epochs for model three to match the performance of the model one and two.\n\n>We can train the model three for even more number of epochs to achieve the accuracy close to 99% on the training dataset."},{"metadata":{"_uuid":"ea2e6dfb5c7bdd29afab863c56be1d17ebd960c4"},"cell_type":"markdown","source":"**USING MOBILENET ARCHITECTURE WITH IMAGE DATA GENERATOR**"},{"metadata":{"trusted":true,"_uuid":"4ca7474ea1854ec5762303a84074d4fd080e557d"},"cell_type":"code","source":"import keras\nfrom keras.applications.mobilenet import MobileNet\nfrom keras.applications.vgg19 import VGG19\nfrom keras.optimizers import Adam\nfrom keras.preprocessing.image import ImageDataGenerator","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"a450606ac0cac8423fdffd6224577afd682fee1f"},"cell_type":"code","source":"datagen = ImageDataGenerator(rescale = 1 / 255.,\n                            horizontal_flip = True,\n                            rotation_range = 10,\n                            width_shift_range = 0.1,\n                            height_shift_range = 0.1)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"63a55364fc37faeed6e2838d19c6bc2f73bc0fce"},"cell_type":"code","source":"mobilenet_model = MobileNet(weights = None, input_shape = (100, 100, 3), classes = 5005)\nmobilenet_model.summary()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"67ce9c8fa1137d616ea37e0f91d8c32123f4b879"},"cell_type":"code","source":"#compile the model\noptim = Adam(lr = 0.001)\nmobilenet_model.compile(loss = 'categorical_crossentropy', optimizer = optim, metrics = ['accuracy'])","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"c50116ee04b2de5a23acffc4ffa6879573f3ac67"},"cell_type":"code","source":"#fit the model on our data\nh2 = mobilenet_model.fit_generator(datagen.flow(x_train, y_train, batch_size = 64), epochs = 300, steps_per_epoch = len(x_train) // 64)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"9c860367d85234fbef40df7ee65ec33476171548"},"cell_type":"code","source":"#let's look how our model performed by plotting the accuracy and loss curves\nsns.set(style = 'darkgrid')\nplt.figure(figsize = (18, 14))\nplt.subplot(2, 1, 1)\nplt.plot(range(300), h2.history['acc'])\nplt.xlabel('EPOCHS')\nplt.ylabel('TRAINING ACCURACY')\nplt.title('TRAINING ACCURACY vs EPOCHS')\n\nplt.figure(figsize = (18, 14))\nplt.subplot(2, 1, 1)\nplt.plot(range(300), h2.history['loss'])\nplt.xlabel('EPOCHS')\nplt.ylabel('TRAINING LOSS')\nplt.title('TRAINING LOSS vs EPOCHS')","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"6051db0154d167e214a149d80bbc9d5ed6f6525d"},"cell_type":"markdown","source":"From the above plot we can observe that after about 90 epochs the loss became stagnant. "},{"metadata":{"_uuid":"f3b0677ddaf41c887b36166fdd4d1d80234d0509"},"cell_type":"markdown","source":" **MAKING PREDICTIONS ON TEST DATA**"},{"metadata":{"trusted":true,"_uuid":"fb1705cc3de2add132eadde861f763d69efb100d"},"cell_type":"code","source":"test_data = os.listdir(\"../input/test/\")\nprint(len(test_data))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"34ffb725b1e2a4e1d247e0ccbc3114e9e5889e61"},"cell_type":"code","source":"test_data = pd.DataFrame(test_data, columns = ['Image'])\ntest_data['Id'] = ''","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"f5275dc5f478682e906278168c49e53ae2e959e6"},"cell_type":"code","source":"x_test = PrepareTrainImages(test_data, test_data.shape[0], \"test\")\nx_test = x_test.astype('float32') / 255","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"3cf84961930da09533f7f48c239ff6b97d81fc22"},"cell_type":"code","source":"predictions = mobilenet_model.predict(np.array(x_test), verbose = 1)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"89085aa02f2f52a73158fd7185c6598b8d325df0"},"cell_type":"code","source":"for i, pred in enumerate(predictions):\n    test_data.loc[i, 'Id'] = ' '.join(le.inverse_transform(pred.argsort()[-5:][::-1]))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"9adf6c7e0eda2cde170beab9bdcbd948a9e7a76d"},"cell_type":"code","source":"test_data.to_csv('model_submission4.csv', index = False)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"023e7b10fa8242980a04c5586eee84134f98805c"},"cell_type":"code","source":"","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.6.6","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"}},"nbformat":4,"nbformat_minor":1}