{"cells":[{"metadata":{"_uuid":"ac29c6450eb3e577474781a126dadc48abd76eff"},"cell_type":"markdown","source":"# **Introduction**\n\nAfter centuries of intense whaling, recovering whale populations still have a hard time adapting to warming oceans and struggle to compete every day with the industrial fishing industry for food.\n\nTo aid whale conservation efforts, scientists use photo surveillance systems to monitor ocean activity. They use the shape of whales’ tails and unique markings found in footage to identify what species of whale they’re analyzing and meticulously log whale pod dynamics and movements. For the past 40 years, most of this work has been done manually by individual scientists, leaving a huge trove of data untapped and underutilized.\n\nIn this competition, we’re challenged to build an algorithm to identify individual whales in images. we’ll analyze Happywhale’s database of over 25,000 images, gathered from research institutions and public contributors. By contributing, we’ll help to open rich fields of understanding for marine mammal population dynamics around the globe."},{"metadata":{"_uuid":"4d7de0db945f943c58b3000b19ce4d62be697356"},"cell_type":"markdown","source":"# **Available Data**\n\nThis training data contains thousands of images of humpback whale flukes. Individual whales have been identified by researchers and given an Id. The challenge is to predict the whale Id of images in the test set. What makes this such a challenge is that there are only a few examples for each of 3,000+ whale Ids.\n\n# File descriptions\n\n* **train.zip** - a folder containing the training images\n* **train.csv** - maps the training Image to the appropriate whale Id. Whales that are not predicted to have a label identified in the training data should be labeled as new_whale.\n* **test.zip** - a folder containing the test images to predict the whale Id\n"},{"metadata":{"_uuid":"9c7095dad3a7d207e4c39690453ff7c6e059accb"},"cell_type":"markdown","source":"# Part 1  - Keras Pre Processing"},{"metadata":{"_uuid":"40e622cf1077880626c68ab52ade566144da50cd","trusted":false},"cell_type":"code","source":"import numpy as np\nimport pandas as pd\nfrom PIL import Image\nimport matplotlib.pyplot as plt\nfrom matplotlib.pyplot import imshow\nfrom IPython.display import HTML\nimport os\nprint(os.listdir(\"../input\"))\n\n%matplotlib inline\n\ndf=pd.read_csv('../input/train.csv')\ndf.head()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"596bd11f5f6259b801a62cc664eed876237a0a62"},"cell_type":"markdown","source":"In the csv file, the feature 'Image' represents the file name of each photos in the train.zip. The feature 'Id' represents the category of the whale in the correspond row feature 'Image. Those whale in the image that doesn't have a label isto be represented as a new_whale. "},{"metadata":{"_uuid":"bfab3fcecd67527e19502e5b12c9208158aafd43","trusted":false},"cell_type":"code","source":"df.count()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"34d0982f7c19dc5c158052c9523563edfd571f34"},"cell_type":"markdown","source":"Let's add a new column into the data frame which indicates the path of each file."},{"metadata":{"_uuid":"d4bc6db2b9a631846b86adac5ff02830fd05a82b","trusted":false},"cell_type":"code","source":"df['Path']=df['Image'].map(lambda x:'../input/train/{}'.format(x))\ndf.head()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"ed5be0ff0b72b1364305e4590e833b8de1f1a403"},"cell_type":"markdown","source":"The feature 'Id' is categorical since it is the label for each Image. It represents the category/species in which each whale in the train data belongs. Since machine learning models need numerical data for processing, we have toencode the categorical content into numerical values. "},{"metadata":{"_uuid":"20cbec543d44f99847c25d75fe9bd3655590fb85","trusted":false},"cell_type":"code","source":"df['Id'].nunique()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"f91f37434b6ebff907bb41e174672a7a09ee0ec0","trusted":false},"cell_type":"code","source":"df['Id'].value_counts().head(20)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"340208ad9424359545dc9b1440dff6ccdf267759"},"cell_type":"markdown","source":"Let's open 2 random whale Image fron the train data. "},{"metadata":{"_uuid":"c7733267d3c999ba5509edb4f8409e440fbc559b","trusted":false},"cell_type":"code","source":"random_whale=np.random.choice(df['Path'],2)\nfor whale in random_whale:\n    image=Image.open(whale)\n    plt.imshow(image)\n    plt.show()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"a4f8fc074b7e118158b53bd4565c89a512ccaf82"},"cell_type":"markdown","source":"Now let's prepare the data for Keras CNN. Let's create x_train and y_train which will be fitted to the keras for training the model. x_train will contain all the images in train dataset and y_train will contain the corresponding Id/label of each whale image.  img_to_array converts a PIL image instance to numpy array. The images will all be reshaped. "},{"metadata":{"_uuid":"b4eefd6e8b08029fc73a6fcf96e1957a264aa49d","trusted":false},"cell_type":"code","source":"from keras.preprocessing import image\nfrom keras.applications.imagenet_utils import preprocess_input\n\ndef add_img(dataset,shape,img_size):\n    \n    x_train = np.zeros((shape, img_size[0], img_size[1], img_size[2]))\n    count = 0\n    \n    for fig in dataset.itertuples():\n        \n        #load train data images into images of specified size\n        img = image.load_img(fig.Path, target_size=img_size)\n        x = image.img_to_array(img)\n        x = preprocess_input(x)\n        x_train[count] = x\n        count += 1\n    \n    return x_train","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"999b8a0ca38de6540b3534f300d14e45b843de55"},"cell_type":"markdown","source":"Now let's prepare the y_train. y_train contain the labels, whale name of each whale image in x_train. The labels are categorical values and hence they need to be numrically encoded. For this LabelEncoder,OneHotEncoder functions from the ScikitLearn library is used. "},{"metadata":{"_uuid":"08cd00d133aa608dd5af1538b36a6fc702f59e45","trusted":false},"cell_type":"code","source":"from sklearn.preprocessing import LabelEncoder\nfrom keras.utils.np_utils import to_categorical\ndef label(y):\n    y_train=np.array(y)\n    label_encoder = LabelEncoder()\n    y_train = label_encoder.fit_transform(y_train)\n    y_train = to_categorical(y_train, num_classes = 5005)\n    return y_train,label_encoder","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"3dca2b79c68f960296565c5cfbe951194b32f527","trusted":false},"cell_type":"code","source":"x_train=add_img(df,df.shape[0],(100,100,3))\ny_train,encoder=label(df['Id'])\nx_train/=255 #Normalizing the data","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"3e6029a1b8361d0e948db37c9560ac8b2b9aba6d","trusted":false},"cell_type":"code","source":"y_train.shape","execution_count":null,"outputs":[]},{"metadata":{"_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","trusted":false},"cell_type":"code","source":"# Importing the Keras packages\nfrom keras.models import Sequential\nfrom keras.layers import Convolution2D\nfrom keras.layers import MaxPooling2D\nfrom keras.layers import Flatten\nfrom keras.layers import Dense\nfrom keras.layers import Dropout\nfrom keras.layers.normalization import BatchNormalization\nfrom keras.preprocessing.image import ImageDataGenerator","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"c26579a89007c37e1727f2a99e1005fb85dc0efb"},"cell_type":"markdown","source":"# Initialising the CNN"},{"metadata":{"_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","trusted":false},"cell_type":"code","source":"classifier = Sequential()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"a3dbdf0119e3fb06417b8bc00ca174d0ce181c5c"},"cell_type":"markdown","source":"# Step 1 - Convolution"},{"metadata":{"_uuid":"45e3e5bef08e9836e256bc761228000ef2dbde88","trusted":false},"cell_type":"code","source":"classifier.add(Convolution2D(16, 5, 5, input_shape = (100,100, 3), activation = 'relu'))","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"273ea78053addba4d909161315a1bcf40d120097"},"cell_type":"markdown","source":"# Step 2 - Pooling"},{"metadata":{"_uuid":"88a4e6e1a2c75f16ae4b1276c65aa1d89b9c5a3c","trusted":false},"cell_type":"code","source":"classifier.add(MaxPooling2D(pool_size = (2, 2)))","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"35b368d1e15d81ac52fcb01dc038e0b6f4c2d88f"},"cell_type":"markdown","source":"# Adding a second convolutional layer"},{"metadata":{"_uuid":"a9f78335cc57f9631160e7dd6c90e50d5cd0d621"},"cell_type":"markdown","source":"Convolution is the first layer to extract features from an input image. Convolution preserves the relationship between pixels by learning image features using small squares of input data. It is a mathematical operation that takes two inputs such as image matrix and a filter or kernal\n\nMax pooling is a type of operation that is typically added to CNNs following individual convolutional layers. When added to a model, max pooling reduces the dimensionality of images by reducing the number of pixels in the output from the previous convolutional layer.\n\nDropout is a technique where randomly selected neurons are ignored during training. They are “dropped-out” randomly. This means that their contribution to the activation of downstream neurons is temporally removed on the forward pass and any weight updates are not applied to the neuron on the backward pass. As a neural network learns, neuron weights settle into their context within the network. Weights of neurons are tuned for specific features providing some specialization. Neighboring neurons become to rely on this specialization, which if taken too far can result in a fragile model too specialized to the training data. This reliant on context for a neuron during training is referred to complex co-adaptations. You can imagine that if neurons are randomly dropped out of the network during training, that other neurons will have to step in and handle the representation required to make predictions for the missing neurons. This is believed to result in multiple independent internal representations being learned by the network. The effect is that the network becomes less sensitive to the specific weights of neurons. This in turn results in a network that is capable of better generalization and is less likely to overfit the training data.\n\nFlatten() flattens the output and feed into a fully connected layer (FC Layer)"},{"metadata":{"_uuid":"a81a701629b672c87f9e2a78e71793ea3c55a155","trusted":false},"cell_type":"code","source":"classifier.add(Convolution2D(16, 5, 5, activation = 'relu'))\nclassifier.add(MaxPooling2D(pool_size = (2, 2)))\nclassifier.add(Dropout(0.25))","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"6024cd4967ee2450f3100a6c8747a4b9b2c47403","trusted":false},"cell_type":"code","source":"classifier.add(Convolution2D(32, 3, 3, activation = 'relu'))\nclassifier.add(MaxPooling2D(pool_size = (2, 2)))\nclassifier.add(Dropout(0.25))","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"9adf27ed0624d675fd9592961e7a0d26ba8cd063","trusted":false},"cell_type":"code","source":"classifier.add(Convolution2D(64, 3, 3, activation = 'relu'))\nclassifier.add(MaxPooling2D(pool_size = (2, 2)))\nclassifier.add(Dropout(0.25))","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"b104b6a6fc16a94a085a3e32adfdef351444a0ce"},"cell_type":"markdown","source":"# Step 3 - Flattening"},{"metadata":{"_uuid":"51cbff380a91196295460dfa4cdaba2ded0caa07","trusted":false},"cell_type":"code","source":"classifier.add(Flatten())","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"1481b242f733af63e914df8a993e862518aabd2d"},"cell_type":"markdown","source":"# Step 4 - Full connection"},{"metadata":{"_uuid":"ee7889a1d2375e8d3b26c6ab6c86cead4e8dec68","trusted":false},"cell_type":"code","source":"classifier.add(Dense(output_dim = 240, activation = 'relu'))\nclassifier.add(BatchNormalization())\nclassifier.add(Dense(output_dim = y_train.shape[1], activation = 'sigmoid'))","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"fa4b75a3bcdb1e055d43adc0045a1517fe065bcd"},"cell_type":"markdown","source":"# Optimizer And Annealer"},{"metadata":{"_uuid":"6955df55d74cf62d0d29a0f5fe413f8f6cbfc5a8"},"cell_type":"markdown","source":"The Adam optimization algorithm is an extension to stochastic gradient descent that has recently seen broader adoption for deep learning applications in computer vision and natural language processing.\n\nStochastic gradient descent maintains a single learning rate (termed alpha) for all weight updates and the learning rate does not change during training.\n\n**Adaptive Gradient Algorithm** (AdaGrad) that maintains a per-parameter learning rate that improves performance on problems with sparse gradients (e.g. natural language and computer vision problems).\n\n**Root Mean Square Propagation** (RMSProp) that also maintains per-parameter learning rates that are adapted based on the average of recent magnitudes of the gradients for the weight (e.g. how quickly it is changing). This means the algorithm does well on online and non-stationary problems"},{"metadata":{"_uuid":"63bf1f103a53d23953368d089033d1afab893f17","trusted":false},"cell_type":"code","source":"from keras.optimizers import Adam\nfrom keras.callbacks import ReduceLROnPlateau\n\n# Define the optimizer\nadam_optimizer = Adam(lr = 0.001, beta_1 = 0.9, beta_2 = 0.999)\n\n# Set a learning rate annealer\nlearning_rate = ReduceLROnPlateau(monitor='val_acc', \n                                            patience=3, \n                                            verbose=1, \n                                            factor=0.5, \n                                            min_lr=0.00001)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"b3b169d8911a00693d253285982b8768bcd91f94"},"cell_type":"markdown","source":"# Compiling the CNN"},{"metadata":{"_uuid":"c09ba43d870488c3bc334ff689b6703965264df5","trusted":false},"cell_type":"code","source":"classifier.compile(optimizer = adam_optimizer, loss = 'categorical_crossentropy', metrics = ['accuracy'])","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"0e7e428dcba909aa86c24a1f997f2cc70e95862a","trusted":false},"cell_type":"code","source":"classifier.summary()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"1a4d4adde871539ffe280c0f57af7b14fbdc1e0d"},"cell_type":"markdown","source":"# Part 2 - Fitting the CNN to the images"},{"metadata":{"_uuid":"61de8cec65b5be73d3815fdea93189c148496e34","trusted":false},"cell_type":"code","source":"whale_detector = classifier.fit(x_train, y_train, epochs=60, batch_size=1000, verbose=10, callbacks=[learning_rate])","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"e46c75c2531bc7f2d13b6be5d8fc52dc932d2f7d"},"cell_type":"markdown","source":"# Let's Predict the model for test Images"},{"metadata":{"_uuid":"415166a30f99b84fa6bf81a99072cd5f0220636f"},"cell_type":"markdown","source":"The method listdir() returns a list containing the names of the entries in the directory given by path. The list is made a data frame and the rest is same as we did for the train data. "},{"metadata":{"_uuid":"8529092131092fa1c904b9c6084430a4afd37059","trusted":false},"cell_type":"code","source":"test = os.listdir(\"../input/test/\")\ntest_df = pd.DataFrame(test, columns=['Image'])\ntest_df['Path']=test_df['Image'].map(lambda x:'../input/test/{}'.format(x))\nx_test=add_img(test_df,test_df.shape[0],(100,100,3))\nx_test/255\npred=classifier.predict(np.array(x_test),verbose=1)#Since numpy array is faster than df","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"2aee2d8f2bbb8e30b708eb170e37081c5aef4cc4"},"cell_type":"markdown","source":"# Lets Plot the Loss and Accuracy changers per epoch"},{"metadata":{"_uuid":"60d47b6e2ab6abc3c2c971eefe679b2e9a164b2a","trusted":false},"cell_type":"code","source":"# Plot the loss curve for training\nplt.plot(whale_detector.history['loss'], color='r', label=\"Train Loss\")\nplt.title(\"Train Loss\")\nplt.xlabel(\"Number of Epochs\")\nplt.ylabel(\"Loss\")\nplt.legend()\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"aa1ef359940a4b5497bda11f081eb19e3bdf36c2","trusted":false},"cell_type":"code","source":"# Plot the accuracy curve for training\nplt.plot(whale_detector.history['acc'], color='g', label=\"Train Accuracy\")\nplt.title(\"Train Accuracy\")\nplt.xlabel(\"Number of Epochs\")\nplt.ylabel(\"Accuracy\")\nplt.legend()\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"9666a4fd5494e8982acd42e5decd95bee95b4831"},"cell_type":"markdown","source":"# Submission"},{"metadata":{"_uuid":"810db338144241f22dc0d291a412a1f860c51b5e","trusted":false},"cell_type":"code","source":"test_df['Id']=''\nfor index,prediction in enumerate(pred):\n    test_df.loc[index, 'Id'] = ' '.join(encoder.inverse_transform(prediction.argsort()[-5:][::-1]))\ntest_df.drop(['Path'],axis=1,inplace=True)\ntest_df.to_csv('submission.csv', index=False)\ntest_df.head()","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"codemirror_mode":{"name":"ipython","version":3},"file_extension":".py","mimetype":"text/x-python","name":"python","nbconvert_exporter":"python","pygments_lexer":"ipython3","version":"3.6.6"}},"nbformat":4,"nbformat_minor":1}