{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Intro\nWelcome to the [Happywhale - Whale and Dolphin Identification](https://www.kaggle.com/c/happy-whale-and-dolphin/data) compedition.\n\n![](https://storage.googleapis.com/kaggle-competitions/kaggle/22962/logos/header.png)\n\nIn this competition, the task was to predict individual humpback whales from images of their flukes. Whales and dolphins in this dataset can be identified by shapes, features and markings of dorsal fins, backs, heads and flanks.\n\n**Table of content:**\n1. [Exploratory Data Analysis](#EDA)\n2. [Load Single Image](#LoadSingleImage)\n3. [Plot Examples](#PlotExamples)\n4. [Image Preprocessing](#ImagePreprocessing)\n5. [Data Generator](#DataGenerator)\n6. [Model](#Model)\n\n<font size=\"4\"><span style=\"color: royalblue;\">Please vote the notebook up if it helps you. Feel free to leave a comment above the notebook. Thank you. </span></font>","metadata":{}},{"cell_type":"markdown","source":"# Libraries","metadata":{}},{"cell_type":"code","source":"import os\nimport pandas as pd\nimport numpy as np\nimport cv2\nimport matplotlib.pyplot as plt\n\nfrom sklearn.model_selection import train_test_split\n\nfrom tensorflow.keras.utils import to_categorical, Sequence\nfrom keras.models import Sequential\nfrom keras.layers import Dense, Dropout, Flatten\nfrom tensorflow.keras.optimizers import RMSprop,Adam\nfrom tensorflow.keras.applications import ResNet50","metadata":{"execution":{"iopub.status.busy":"2022-02-10T19:22:17.842775Z","iopub.execute_input":"2022-02-10T19:22:17.843508Z","iopub.status.idle":"2022-02-10T19:22:25.696321Z","shell.execute_reply.started":"2022-02-10T19:22:17.843373Z","shell.execute_reply":"2022-02-10T19:22:25.695299Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Path","metadata":{}},{"cell_type":"code","source":"path = '/kaggle/input/happy-whale-and-dolphin/'\nos.listdir(path)","metadata":{"execution":{"iopub.status.busy":"2022-02-10T19:22:25.697983Z","iopub.execute_input":"2022-02-10T19:22:25.698228Z","iopub.status.idle":"2022-02-10T19:22:25.711755Z","shell.execute_reply.started":"2022-02-10T19:22:25.698199Z","shell.execute_reply":"2022-02-10T19:22:25.710969Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Load Data","metadata":{}},{"cell_type":"code","source":"train_data = pd.read_csv(path+'train.csv')\nsamp_subm = pd.read_csv(path+'sample_submission.csv')","metadata":{"execution":{"iopub.status.busy":"2022-02-10T19:35:03.053976Z","iopub.execute_input":"2022-02-10T19:35:03.054615Z","iopub.status.idle":"2022-02-10T19:35:03.171344Z","shell.execute_reply.started":"2022-02-10T19:35:03.054578Z","shell.execute_reply":"2022-02-10T19:35:03.170150Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"samp_subm.loc[0, 'predictions']","metadata":{"execution":{"iopub.status.busy":"2022-02-10T19:35:21.254876Z","iopub.execute_input":"2022-02-10T19:35:21.256006Z","iopub.status.idle":"2022-02-10T19:35:21.263869Z","shell.execute_reply.started":"2022-02-10T19:35:21.255961Z","shell.execute_reply":"2022-02-10T19:35:21.263186Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Overview","metadata":{}},{"cell_type":"code","source":"print('Number train samples:', len(train_data))\nprint('Number train images:', len(os.listdir(path+'train_images/')))\nprint('Number test images:', len(os.listdir(path+'test_images/')))","metadata":{"execution":{"iopub.status.busy":"2022-02-10T19:22:25.895880Z","iopub.execute_input":"2022-02-10T19:22:25.896130Z","iopub.status.idle":"2022-02-10T19:22:26.417003Z","shell.execute_reply.started":"2022-02-10T19:22:25.896103Z","shell.execute_reply":"2022-02-10T19:22:26.416117Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_data.head()","metadata":{"execution":{"iopub.status.busy":"2022-02-10T19:22:26.419286Z","iopub.execute_input":"2022-02-10T19:22:26.419717Z","iopub.status.idle":"2022-02-10T19:22:26.436804Z","shell.execute_reply.started":"2022-02-10T19:22:26.419684Z","shell.execute_reply":"2022-02-10T19:22:26.435919Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Exploratory Data Analysis <a name=\"EDA\"></a>","metadata":{}},{"cell_type":"markdown","source":"There are 30 different species collected from 28 different research organizations:","metadata":{}},{"cell_type":"code","source":"train_data['species'].value_counts()","metadata":{"execution":{"iopub.status.busy":"2022-02-10T19:22:26.437805Z","iopub.execute_input":"2022-02-10T19:22:26.438413Z","iopub.status.idle":"2022-02-10T19:22:26.459655Z","shell.execute_reply.started":"2022-02-10T19:22:26.438375Z","shell.execute_reply":"2022-02-10T19:22:26.458597Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"There are duplicate names in species which can be merged:\n* bottlenose_dolphin and bottlenose_dolhin,\n* killer_whale and kiler_whale.\n\nSo in total we have 28 different species.","metadata":{}},{"cell_type":"markdown","source":"Individuals have been manually identified and given an individual_id by marine researches, and our task is to correctly identify these individuals in images:","metadata":{}},{"cell_type":"code","source":"train_data['individual_id'].value_counts()","metadata":{"execution":{"iopub.status.busy":"2022-02-10T19:22:26.460857Z","iopub.execute_input":"2022-02-10T19:22:26.461369Z","iopub.status.idle":"2022-02-10T19:22:26.484921Z","shell.execute_reply.started":"2022-02-10T19:22:26.461332Z","shell.execute_reply":"2022-02-10T19:22:26.484047Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Load Single Image <a name=\"LoadSingleImage\"></a>\nWe plot the first image of of the train data.","metadata":{}},{"cell_type":"code","source":"row = 0\nfile = train_data.loc[row, 'image']\nspecies = train_data.loc[row, 'species']\n\nimg = cv2.imread(path+'train_images/'+file)\nprint('Shape:', img.shape)","metadata":{"execution":{"iopub.status.busy":"2022-02-10T19:22:26.486716Z","iopub.execute_input":"2022-02-10T19:22:26.487385Z","iopub.status.idle":"2022-02-10T19:22:26.529183Z","shell.execute_reply.started":"2022-02-10T19:22:26.487338Z","shell.execute_reply":"2022-02-10T19:22:26.528464Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig, ax = plt.subplots(1, 1, figsize=(7, 7))\nax.imshow(cv2.cvtColor(img, cv2.COLOR_BGR2RGB))\nax.set_xticklabels([])\nax.set_yticklabels([])\nax.set_title(species)\nplt.show()","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-02-10T19:22:26.530289Z","iopub.execute_input":"2022-02-10T19:22:26.530677Z","iopub.status.idle":"2022-02-10T19:22:26.960094Z","shell.execute_reply.started":"2022-02-10T19:22:26.530646Z","shell.execute_reply":"2022-02-10T19:22:26.958982Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Plot Examples <a name=\"PlotExamples\"></a>\nWe plot example images of the species top 3.","metadata":{}},{"cell_type":"code","source":"def plot_examples(category = 'bottlenose_dolphin'):\n    \"\"\" Plot 5 images of a given category \"\"\"\n    \n    fig, axs = plt.subplots(1, 5, figsize=(25, 20))\n    fig.subplots_adjust(hspace = .1, wspace=.1)\n    axs = axs.ravel()\n    temp = train_data[train_data['species']==category].copy()\n    temp.index = range(len(temp.index))\n    for i in range(5):\n        file = temp.loc[i, 'image']\n        species = temp.loc[i, 'species']\n        img = cv2.imread(path+'train_images/'+file)\n        axs[i].imshow(cv2.cvtColor(img, cv2.COLOR_BGR2RGB))\n        axs[i].set_title(species)\n        axs[i].set_xticklabels([])\n        axs[i].set_yticklabels([])\n    plt.show()","metadata":{"execution":{"iopub.status.busy":"2022-02-10T19:22:26.961502Z","iopub.execute_input":"2022-02-10T19:22:26.962307Z","iopub.status.idle":"2022-02-10T19:22:26.971143Z","shell.execute_reply.started":"2022-02-10T19:22:26.962258Z","shell.execute_reply":"2022-02-10T19:22:26.970305Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plot_examples(category = 'bottlenose_dolphin')","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-02-10T19:22:26.972291Z","iopub.execute_input":"2022-02-10T19:22:26.972880Z","iopub.status.idle":"2022-02-10T19:22:32.289792Z","shell.execute_reply.started":"2022-02-10T19:22:26.972846Z","shell.execute_reply":"2022-02-10T19:22:32.288692Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plot_examples(category = 'beluga')","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-02-10T19:22:32.291564Z","iopub.execute_input":"2022-02-10T19:22:32.291871Z","iopub.status.idle":"2022-02-10T19:22:33.732398Z","shell.execute_reply.started":"2022-02-10T19:22:32.291834Z","shell.execute_reply":"2022-02-10T19:22:33.731545Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plot_examples(category = 'humpback_whale')","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-02-10T19:22:33.734157Z","iopub.execute_input":"2022-02-10T19:22:33.734465Z","iopub.status.idle":"2022-02-10T19:22:38.519451Z","shell.execute_reply.started":"2022-02-10T19:22:33.734428Z","shell.execute_reply":"2022-02-10T19:22:38.518591Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Image Preprocessing <a name=\"ImagePreprocessing\"></a>\nAs we can see the images have different format: landscape or portrait. For the neural network we need a standard size. So we have to prepare the data. ","metadata":{}},{"cell_type":"code","source":"def image_preprocessing(image, image_size):\n    \"\"\" Image Preprocessing \"\"\"\n    \n    # Crop Image\n    mid_row = int(image.shape[0]/2)\n    mid_col = int(image.shape[1]/2)\n    if image.shape[0]>image.shape[1]:\n        image_cropped = image[mid_row-mid_col:mid_row+mid_col,\n                                   0:image.shape[1]]\n    else:\n        image_cropped = image[0:image.shape[0],\n                                   mid_col-mid_row:mid_col+mid_row]\n    \n    # Rescale Image\n    image_rescale = cv2.resize(image_cropped,\n                               dsize=(image_size, image_size))\n    return image_rescale\n\n\ndef plot_befor_after(image):\n    \"\"\" Compare original and prepared image \"\"\"\n    \n    fig, axs = plt.subplots(1, 2, figsize=(15, 10))\n    fig.subplots_adjust(hspace = .1, wspace=.1)\n    axs = axs.ravel()\n    # Plot Original Image\n    axs[0].imshow(cv2.cvtColor(image, cv2.COLOR_BGR2RGB))\n    axs[0].set_title('original shape: '+str(image.shape))\n    # Image Preprocessing\n    image_rescale = image_preprocessing(image, image_size)\n    # Plot Prepared Image\n    axs[1].imshow(cv2.cvtColor(image_rescale, cv2.COLOR_BGR2RGB))\n    axs[1].set_title('rescaled shape: '+str(image_rescale.shape))\n    for i in range(2):\n        axs[i].set_xticklabels([])\n        axs[i].set_yticklabels([])\n    plt.show()","metadata":{"execution":{"iopub.status.busy":"2022-02-10T19:22:38.523837Z","iopub.execute_input":"2022-02-10T19:22:38.524196Z","iopub.status.idle":"2022-02-10T19:22:38.540277Z","shell.execute_reply.started":"2022-02-10T19:22:38.524152Z","shell.execute_reply":"2022-02-10T19:22:38.539171Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"image_size = 128","metadata":{"execution":{"iopub.status.busy":"2022-02-10T19:22:38.541461Z","iopub.execute_input":"2022-02-10T19:22:38.542065Z","iopub.status.idle":"2022-02-10T19:22:38.555716Z","shell.execute_reply.started":"2022-02-10T19:22:38.542027Z","shell.execute_reply":"2022-02-10T19:22:38.554954Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"row = 2022\nfile = train_data.loc[row, 'image']\nspecies = train_data.loc[row, 'species']\nimage = cv2.imread(path+'train_images/'+file)\nprint('Shape:', image.shape)","metadata":{"execution":{"iopub.status.busy":"2022-02-10T19:22:38.557161Z","iopub.execute_input":"2022-02-10T19:22:38.557664Z","iopub.status.idle":"2022-02-10T19:22:38.612210Z","shell.execute_reply.started":"2022-02-10T19:22:38.557610Z","shell.execute_reply":"2022-02-10T19:22:38.611540Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plot_befor_after(image)","metadata":{"execution":{"iopub.status.busy":"2022-02-10T19:22:38.613263Z","iopub.execute_input":"2022-02-10T19:22:38.613608Z","iopub.status.idle":"2022-02-10T19:22:39.181213Z","shell.execute_reply.started":"2022-02-10T19:22:38.613579Z","shell.execute_reply":"2022-02-10T19:22:39.180274Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Split Data","metadata":{}},{"cell_type":"code","source":"list_IDs_train, list_IDs_val = train_test_split(list(train_data.index), test_size=0.33, random_state=2022)\nlist_IDs_test = list(samp_subm.index)","metadata":{"execution":{"iopub.status.busy":"2022-02-10T19:22:39.182784Z","iopub.execute_input":"2022-02-10T19:22:39.183172Z","iopub.status.idle":"2022-02-10T19:22:39.214010Z","shell.execute_reply.started":"2022-02-10T19:22:39.183113Z","shell.execute_reply":"2022-02-10T19:22:39.213035Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print('Number train samples:', len(list_IDs_train))\nprint('Number val samples:', len(list_IDs_val))\nprint('Number test samples:', len(list_IDs_test))","metadata":{"execution":{"iopub.status.busy":"2022-02-10T19:22:39.215284Z","iopub.execute_input":"2022-02-10T19:22:39.215911Z","iopub.status.idle":"2022-02-10T19:22:39.221406Z","shell.execute_reply.started":"2022-02-10T19:22:39.215848Z","shell.execute_reply":"2022-02-10T19:22:39.220756Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Data Generator <a name=\"DataGenerator\"></a>\nWe define a data generator to load the data on demand.","metadata":{}},{"cell_type":"code","source":"img_size = 32\nimg_channel = 3\nbatch_size = 64\nnum_classes = len(train_data['species'].value_counts())","metadata":{"execution":{"iopub.status.busy":"2022-02-10T19:22:39.222418Z","iopub.execute_input":"2022-02-10T19:22:39.223201Z","iopub.status.idle":"2022-02-10T19:22:39.240452Z","shell.execute_reply.started":"2022-02-10T19:22:39.223150Z","shell.execute_reply":"2022-02-10T19:22:39.239444Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"class DataGenerator(Sequence):\n    def __init__(self, path, list_IDs, data, img_size, img_channel, batch_size, num_classes):\n        self.path = path\n        self.list_IDs = list_IDs\n        self.data = data\n        self.img_size = img_size\n        self.img_channel = img_channel\n        self.batch_size = batch_size\n        self.num_classes = num_classes\n        self.indexes = np.arange(len(self.list_IDs))\n        if self.path.find('train')>=0:\n            self.labels = pd.get_dummies(self.data['species'])\n        \n    def __len__(self):\n        len_ = int(len(self.list_IDs)/self.batch_size)\n        if len_*self.batch_size < len(self.list_IDs):\n            len_ += 1\n        return len_\n    \n    def __getitem__(self, index):\n        indexes = self.indexes[index*self.batch_size:(index+1)*self.batch_size]\n        list_IDs_temp = [self.list_IDs[k] for k in indexes]\n        X, y = self.__data_generation(list_IDs_temp)\n        return X, y\n    \n    def __data_generation(self, list_IDs_temp):\n        X = np.zeros((self.batch_size, self.img_size, self.img_size, self.img_channel))\n        y = np.zeros((self.batch_size, self.num_classes), dtype=int)\n        for i, ID in enumerate(list_IDs_temp):\n            \n            file = self.data.loc[ID, 'image']\n            \n            img = cv2.imread(self.path+file)\n            \n            img_prep = image_preprocessing(img, self.img_size)\n            X[i, ] = img_prep/255\n            if self.path.find('train')>=0:\n                y[i, ] = self.labels.loc[ID]\n        return X, y","metadata":{"execution":{"iopub.status.busy":"2022-02-10T19:22:39.242010Z","iopub.execute_input":"2022-02-10T19:22:39.242518Z","iopub.status.idle":"2022-02-10T19:22:39.259478Z","shell.execute_reply.started":"2022-02-10T19:22:39.242469Z","shell.execute_reply":"2022-02-10T19:22:39.258624Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_generator = DataGenerator(path+'train_images/', list_IDs_train, train_data, img_size, img_channel, batch_size, num_classes)\nval_generator = DataGenerator(path+'train_images/', list_IDs_val, train_data, img_size, img_channel, batch_size, num_classes)\ntest_generator = DataGenerator(path+'test_images/', list_IDs_test, samp_subm, img_size, img_channel, batch_size, num_classes)","metadata":{"execution":{"iopub.status.busy":"2022-02-10T19:22:39.260696Z","iopub.execute_input":"2022-02-10T19:22:39.260947Z","iopub.status.idle":"2022-02-10T19:22:39.307496Z","shell.execute_reply.started":"2022-02-10T19:22:39.260915Z","shell.execute_reply":"2022-02-10T19:22:39.306813Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Test data generator:","metadata":{}},{"cell_type":"code","source":"X, y = train_generator.__getitem__(0)\nX[0].shape","metadata":{"execution":{"iopub.status.busy":"2022-02-10T19:22:39.308650Z","iopub.execute_input":"2022-02-10T19:22:39.309279Z","iopub.status.idle":"2022-02-10T19:22:43.975950Z","shell.execute_reply.started":"2022-02-10T19:22:39.309239Z","shell.execute_reply":"2022-02-10T19:22:43.974978Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Model <a name=\"Model\"></a>\n**Coming soon**","metadata":{}},{"cell_type":"code","source":"weights='../input/models/resnet50_weights_tf_dim_ordering_tf_kernels_notop.h5'\nconv_base = ResNet50(weights=weights,\n                     include_top=False,\n                     input_shape=(img_size, img_size, img_channel))\nconv_base.trainable = True","metadata":{"execution":{"iopub.status.busy":"2022-02-10T19:22:43.977143Z","iopub.execute_input":"2022-02-10T19:22:43.977359Z","iopub.status.idle":"2022-02-10T19:22:47.219676Z","shell.execute_reply.started":"2022-02-10T19:22:43.977333Z","shell.execute_reply":"2022-02-10T19:22:47.218982Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Export","metadata":{}},{"cell_type":"code","source":"samp_subm['predictions'] = '37c7aba965a5 114207cab555 a6e325d8e924 new_individual'\nsamp_subm.head()","metadata":{"execution":{"iopub.status.busy":"2022-02-10T19:35:54.725606Z","iopub.execute_input":"2022-02-10T19:35:54.725959Z","iopub.status.idle":"2022-02-10T19:35:54.738961Z","shell.execute_reply.started":"2022-02-10T19:35:54.725919Z","shell.execute_reply":"2022-02-10T19:35:54.737796Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"samp_subm.to_csv('submission.csv', index=False)","metadata":{"execution":{"iopub.status.busy":"2022-02-10T19:22:47.221243Z","iopub.execute_input":"2022-02-10T19:22:47.221740Z","iopub.status.idle":"2022-02-10T19:22:47.361622Z","shell.execute_reply.started":"2022-02-10T19:22:47.221685Z","shell.execute_reply":"2022-02-10T19:22:47.360966Z"},"trusted":true},"execution_count":null,"outputs":[]}]}