{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# [Happywhale - Whale and Dolphin Identification](https://www.kaggle.com/c/happy-whale-and-dolphin)","metadata":{}},{"cell_type":"markdown","source":"## Notebook Contents\n1. [Introduction](#introduction)\n2. [Submission Format](#submission-format)\n3. [Evaluation Metric Explained](#evaluation-metric-explained)\n4. [Loading Dataset](#loading-dataset)\n5. [Data Cleaning](#data-cleaning)\n6. [Dataset Visualization](#visualization)<br/>\n     6.1 [Visualize Train and Test Images](#visualization)<br/>\n     6.2 [Visualize Class Distribution](#class-distribution-analysis)<br/>\n     6.3 [Observations](#observation-regarding-class-distribution)<br/>\n7. [Getting Image Resolutions](#image-resolutions)\n8. [Color Analysis](#color-analysis)<br/>\n    8.1 [Check Gray Scale Images](#color-analysis)<br/>\n    8.2 [Visualize Mean Intensity for RGB Channels](#get-mean-intensity-for-each-channel-RGB)<br/>\n    8.3 [Observations](#observation-regarding-color-distribution)<br/>\n9. [Data Augmentation](#data-augmentation)\n10. [Preprocessing Dataset](#preprocessing)\n\n<br>\n\n<a id=\"introduction\"></a>\n# Introduction\nThis training data contains thousands of images of whales and dolphins. Individual whales and dolphins have been identified by researchers and given an `Id`. The challenge is to predict the `Id` of images in the test set by unique—but often subtle—characteristics of their natural markings. The best submissions will suggest photo-`Id` solutions that are fast and accurate.\n\n<br>\n\n### If you find this notebook useful,  <font color='red'>please support with an upvote</font> 🙏","metadata":{}},{"cell_type":"markdown","source":"# Importing Libraries","metadata":{}},{"cell_type":"code","source":"!pip install pycaret","metadata":{"execution":{"iopub.status.busy":"2022-04-30T00:15:03.133877Z","iopub.execute_input":"2022-04-30T00:15:03.134590Z","iopub.status.idle":"2022-04-30T00:15:42.734291Z","shell.execute_reply.started":"2022-04-30T00:15:03.134505Z","shell.execute_reply":"2022-04-30T00:15:42.733438Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import os\n\nimport pandas as pd\nimport numpy as np\nimport tensorflow as tf\nimport matplotlib.pyplot as plt\nimport seaborn as sns\nimport pickle\nimport shutil\n\nfrom keras import layers\nfrom keras.models import Sequential\nfrom keras.preprocessing import image\nfrom keras.layers import Input, Dense, Activation, Dropout\nfrom keras.layers import Flatten, BatchNormalization, Conv2D\nfrom keras.layers import MaxPooling2D, AveragePooling2D\nfrom keras.applications.imagenet_utils import preprocess_input\n\nfrom PIL import Image\nfrom tqdm import tqdm\nimport random as rnd\nimport cv2\n\n!pip install livelossplot\nfrom livelossplot import PlotLossesKeras\n\n%matplotlib inline","metadata":{"_kg_hide-output":true,"execution":{"iopub.status.busy":"2022-04-30T03:08:54.005091Z","iopub.execute_input":"2022-04-30T03:08:54.005393Z","iopub.status.idle":"2022-04-30T03:09:08.135759Z","shell.execute_reply.started":"2022-04-30T03:08:54.005360Z","shell.execute_reply":"2022-04-30T03:09:08.134862Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import numpy as np\nimport os\nimport PIL\nimport PIL.Image\nimport tensorflow as tf\nimport tensorflow_datasets as tfds","metadata":{"execution":{"iopub.status.busy":"2022-04-30T00:15:56.321574Z","iopub.execute_input":"2022-04-30T00:15:56.321867Z","iopub.status.idle":"2022-04-30T00:15:57.185289Z","shell.execute_reply.started":"2022-04-30T00:15:56.321826Z","shell.execute_reply":"2022-04-30T00:15:57.184543Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"submission-format\"></a>\n# Submission Format\n\n### We need to predict 5 labels for each of the image.\nFor each image in the test set, we can predict up to 5 individual_id labels. There are individuals in the test set that are not seen in the training data; these should be predicted as new_individual. The file should contain a header and have the following format:\n\n```\nimage,predictions \n000188a72f2562.jpg,37c7aba965a5 114207cab555 a6e325d8e924 19fbb960f07d new_individual \n000ba09273d6f3.jpg,37c7aba965a5 114207cab555 a6e325d8e924 19fbb960f07d new_individual \n...\n```","metadata":{}},{"cell_type":"markdown","source":"<br>\n\n<a id=\"loading-dataset\"></a>\n# Loading Dataset\nWe'll use here the [Pandas](https://pandas.pydata.org/pandas-docs/stable/) to load the dataset into memory","metadata":{}},{"cell_type":"code","source":"train_df = pd.read_csv('../input/happy-whale-and-dolphin/train.csv')\ntrain_df['path'] = '../input/happy-whale-and-dolphin/train_images/' + train_df['image']\n\npred_df = pd.read_csv('../input/happy-whale-and-dolphin/sample_submission.csv')\npred_df['path'] = '../input/happy-whale-and-dolphin/test_images/' + pred_df['image']","metadata":{"execution":{"iopub.status.busy":"2022-04-30T00:15:57.187293Z","iopub.execute_input":"2022-04-30T00:15:57.187487Z","iopub.status.idle":"2022-04-30T00:15:57.329564Z","shell.execute_reply.started":"2022-04-30T00:15:57.187463Z","shell.execute_reply":"2022-04-30T00:15:57.328830Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Having two csv files\n* train.csv - contain image name,species and individual_id\n*  sample_submission.csv - contain image name, dummy label for the images in the test folder\n\n#### And two folders contain the images\n* train - having 51033 images of different type of whales and dolphins. There Labels have provided in the train.csv file\n* test - having 27956 images of different type of whales and dolphins. We need to predict their labels","metadata":{}},{"cell_type":"code","source":"train_df.head(10)","metadata":{"execution":{"iopub.status.busy":"2022-04-30T00:15:57.332280Z","iopub.execute_input":"2022-04-30T00:15:57.332809Z","iopub.status.idle":"2022-04-30T00:15:57.350233Z","shell.execute_reply.started":"2022-04-30T00:15:57.332741Z","shell.execute_reply":"2022-04-30T00:15:57.349591Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print('Train samples count: ', len(train_df))\ntrain_df.columns","metadata":{"execution":{"iopub.status.busy":"2022-04-30T00:15:57.351525Z","iopub.execute_input":"2022-04-30T00:15:57.351789Z","iopub.status.idle":"2022-04-30T00:15:57.359492Z","shell.execute_reply.started":"2022-04-30T00:15:57.351750Z","shell.execute_reply":"2022-04-30T00:15:57.358664Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print('Species Count: ',len(train_df['species'].value_counts()))\ntrain_df['species'].value_counts()","metadata":{"execution":{"iopub.status.busy":"2022-04-30T00:15:57.361129Z","iopub.execute_input":"2022-04-30T00:15:57.361685Z","iopub.status.idle":"2022-04-30T00:15:57.386087Z","shell.execute_reply.started":"2022-04-30T00:15:57.361626Z","shell.execute_reply":"2022-04-30T00:15:57.385400Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"data-cleaning\"></a>\n# Data Cleaning\n### Fixing Duplicate Labels\n* `bottlenose_dolpin` -> `bottlenose_dolphin`\n* `kiler_whale` -> `killer_whale`\n* `beluga` -> `beluga_whale`\n\n### Changing Label due to extreme similarities\n* `globis` & `pilot_whale` -> `short_finned_pilot_whale`","metadata":{}},{"cell_type":"code","source":"print('Before fixing duplicate labels : ')\nprint(\"Number of unique species : \", train_df['species'].nunique())\n\ntrain_df['species'].replace({\n    'bottlenose_dolpin' : 'bottlenose_dolphin',\n    'kiler_whale' : 'killer_whale',\n    'beluga' : 'beluga_whale',\n    'globis' : 'short_finned_pilot_whale',\n    'pilot_whale' : 'short_finned_pilot_whale'\n},inplace =True)\n\nprint('\\nAfter fixing duplicate labels : ')\nprint(\"Number of unique species : \", train_df['species'].nunique())\n\n\ntrain_df['class'] = train_df['species'].apply(lambda x: x.split('_')[-1])\ntrain_df.head()","metadata":{"execution":{"iopub.status.busy":"2022-04-30T00:15:57.387003Z","iopub.execute_input":"2022-04-30T00:15:57.387190Z","iopub.status.idle":"2022-04-30T00:15:57.448205Z","shell.execute_reply.started":"2022-04-30T00:15:57.387166Z","shell.execute_reply":"2022-04-30T00:15:57.447539Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Checking missing data\nLets check if there is any missing values in our dataset","metadata":{}},{"cell_type":"code","source":"train_df.isna().sum()","metadata":{"execution":{"iopub.status.busy":"2022-04-30T00:15:57.449395Z","iopub.execute_input":"2022-04-30T00:15:57.450084Z","iopub.status.idle":"2022-04-30T00:15:57.481795Z","shell.execute_reply.started":"2022-04-30T00:15:57.450046Z","shell.execute_reply":"2022-04-30T00:15:57.480892Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Check for missing image\nNow lets see if there is any missing image","metadata":{}},{"cell_type":"code","source":"len(os.listdir('../input/happy-whale-and-dolphin/train_images')),len(train_df)","metadata":{"execution":{"iopub.status.busy":"2022-04-30T00:15:57.486401Z","iopub.execute_input":"2022-04-30T00:15:57.486621Z","iopub.status.idle":"2022-04-30T00:15:57.958846Z","shell.execute_reply.started":"2022-04-30T00:15:57.486591Z","shell.execute_reply":"2022-04-30T00:15:57.958147Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"visualization\"></a>\n# Visualization\n### Getting all unique species","metadata":{}},{"cell_type":"code","source":"# Getting the photos of all unique species\nplt.figure(figsize = (15,12))\nfor idx,i in enumerate(train_df.species.unique()):\n    plt.subplot(4,7,idx+1)\n    df = train_df[train_df['species'] ==i].reset_index(drop = True)\n    image_path = df.loc[rnd.randint(0, len(df))-1,'path']\n    img = Image.open(image_path)\n    img = img.resize((224,224))\n    plt.imshow(img)\n    plt.axis('off')\n    plt.title(i)\nplt.tight_layout()\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-04-30T00:15:57.960172Z","iopub.execute_input":"2022-04-30T00:15:57.960450Z","iopub.status.idle":"2022-04-30T00:16:02.582151Z","shell.execute_reply.started":"2022-04-30T00:15:57.960414Z","shell.execute_reply":"2022-04-30T00:16:02.580357Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Function to plot whales \ndef plot_species(df,species_name):\n    plt.figure(figsize = (12,12))\n    species_df = df[df['species'] ==species_name].reset_index(drop = True)\n    plt.suptitle(species_name)\n    for idx,i in enumerate(np.random.choice(species_df['path'],5)):\n        plt.subplot(8,8,idx+1)\n        image_path = i\n        img = Image.open(image_path)\n        img = img.resize((224,224))\n        plt.imshow(img)\n        plt.axis('off')\n    plt.tight_layout()\n    plt.show()","metadata":{"execution":{"iopub.status.busy":"2022-04-30T00:16:02.583130Z","iopub.execute_input":"2022-04-30T00:16:02.583356Z","iopub.status.idle":"2022-04-30T00:16:02.591296Z","shell.execute_reply.started":"2022-04-30T00:16:02.583326Z","shell.execute_reply":"2022-04-30T00:16:02.590701Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### more images from each species","metadata":{}},{"cell_type":"code","source":"for species in train_df['species'].unique():\n    #print('\\n\\n')\n    plot_species(train_df , species)","metadata":{"_kg_hide-output":false,"scrolled":true,"execution":{"iopub.status.busy":"2022-04-30T00:16:02.592880Z","iopub.execute_input":"2022-04-30T00:16:02.593324Z","iopub.status.idle":"2022-04-30T00:16:24.584113Z","shell.execute_reply.started":"2022-04-30T00:16:02.593288Z","shell.execute_reply":"2022-04-30T00:16:24.583335Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Lets see some image by individual_id\n\nWe have to predict individual_id from image. So lets see how each individual looks like.","metadata":{}},{"cell_type":"code","source":"def plot_individual(df,individual_id):\n    plt.figure(figsize = (12,12))\n    species_df = df[df['individual_id'] == individual_id].reset_index(drop = True)\n    plt.suptitle(individual_id)\n    for idx,i in enumerate(np.random.choice(species_df['path'],24)):\n        plt.subplot(8,8,idx+1)\n        image_path = i\n        img = Image.open(image_path)\n        img = img.resize((224,224))\n        plt.imshow(img)\n        plt.axis('off')\n    plt.tight_layout()\n    plt.show()","metadata":{"execution":{"iopub.status.busy":"2022-04-30T00:16:24.585470Z","iopub.execute_input":"2022-04-30T00:16:24.586805Z","iopub.status.idle":"2022-04-30T00:16:24.594050Z","shell.execute_reply.started":"2022-04-30T00:16:24.586759Z","shell.execute_reply":"2022-04-30T00:16:24.592963Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Top 5 most frequent individual","metadata":{}},{"cell_type":"code","source":"top_5_ids = train_df.individual_id.value_counts().head(5)\nfor i in top_5_ids.index:\n    #print('\\n\\n')\n    plot_individual(train_df , i)","metadata":{"execution":{"iopub.status.busy":"2022-04-30T00:16:24.596840Z","iopub.execute_input":"2022-04-30T00:16:24.597037Z","iopub.status.idle":"2022-04-30T00:16:52.130101Z","shell.execute_reply.started":"2022-04-30T00:16:24.597013Z","shell.execute_reply":"2022-04-30T00:16:52.129490Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Top 5 least frequent individual\n\nWe will get duplicate images because many individual has only one training image.","metadata":{}},{"cell_type":"code","source":"last_5_ids = train_df.individual_id.value_counts().tail(5)\nfor i in last_5_ids.index:\n    #print('\\n\\n')\n    plot_individual(train_df , i)","metadata":{"execution":{"iopub.status.busy":"2022-04-30T00:16:52.131377Z","iopub.execute_input":"2022-04-30T00:16:52.131776Z","iopub.status.idle":"2022-04-30T00:17:09.028547Z","shell.execute_reply.started":"2022-04-30T00:16:52.131738Z","shell.execute_reply":"2022-04-30T00:17:09.027889Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Seeing the distribution of individuals","metadata":{}},{"cell_type":"code","source":"train_df.individual_id.value_counts().describe()","metadata":{"execution":{"iopub.status.busy":"2022-04-30T00:17:09.029697Z","iopub.execute_input":"2022-04-30T00:17:09.030803Z","iopub.status.idle":"2022-04-30T00:17:09.053908Z","shell.execute_reply.started":"2022-04-30T00:17:09.030761Z","shell.execute_reply":"2022-04-30T00:17:09.052811Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Lets see some test images","metadata":{}},{"cell_type":"code","source":"t_df = pd.read_csv('../input/happy-whale-and-dolphin/sample_submission.csv')\nt_df['path'] = '../input/happy-whale-and-dolphin/test_images/' + t_df['image']\n\ndef plot_testimages(df):\n    plt.figure(figsize = (12,12))\n    plt.suptitle('Test Images')\n    for idx,i in enumerate(np.random.choice(df['path'],48)):\n        plt.subplot(8,8,idx+1)\n        image_path = i\n        img = Image.open(image_path)\n        img = img.resize((224,224))\n        plt.imshow(img)\n        plt.axis('off')\n    plt.tight_layout()\n    plt.show()\n\nplot_testimages(t_df)\ndel t_df","metadata":{"execution":{"iopub.status.busy":"2022-04-30T00:17:09.055487Z","iopub.execute_input":"2022-04-30T00:17:09.055872Z","iopub.status.idle":"2022-04-30T00:17:17.064055Z","shell.execute_reply.started":"2022-04-30T00:17:09.055832Z","shell.execute_reply":"2022-04-30T00:17:17.062614Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Observations regarding handpicked images\n\n1. There are some abnormal images in both train and test dataset\n2. Some training images contains people, boats, birds, penguins etc\n3. Many training images are cropped but some are not.\n4. The uncropped images must be taken care of.\n5. There are some images take from under water","metadata":{}},{"cell_type":"markdown","source":"# Class Distribution Analysis","metadata":{}},{"cell_type":"code","source":"sns.countplot(x='class',data=train_df)\nplt.title('Distribution of classes')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-04-30T00:17:17.065333Z","iopub.execute_input":"2022-04-30T00:17:17.065753Z","iopub.status.idle":"2022-04-30T00:17:17.498271Z","shell.execute_reply.started":"2022-04-30T00:17:17.065720Z","shell.execute_reply":"2022-04-30T00:17:17.497494Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Percentage of images of whale and dolphin in the dataset","metadata":{}},{"cell_type":"code","source":"plt.figure(figsize=(5,5))\nclass_cnt = train_df.groupby(['class']).size().reset_index(name = 'counts')\nplt.pie(class_cnt['counts'], labels=class_cnt['class'],colors=['deepskyblue','royalblue'], autopct='%1.1f%%')\nplt.legend(loc='upper left')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-04-30T00:17:17.502086Z","iopub.execute_input":"2022-04-30T00:17:17.502286Z","iopub.status.idle":"2022-04-30T00:17:17.614478Z","shell.execute_reply.started":"2022-04-30T00:17:17.502259Z","shell.execute_reply":"2022-04-30T00:17:17.613682Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Number of training images of each species","metadata":{}},{"cell_type":"code","source":"plt.figure(figsize=(8,8))\nsns.countplot(data=train_df, y = 'species',  palette='mako', dodge=False)\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-04-30T00:17:17.616208Z","iopub.execute_input":"2022-04-30T00:17:17.616768Z","iopub.status.idle":"2022-04-30T00:17:18.016327Z","shell.execute_reply.started":"2022-04-30T00:17:17.616725Z","shell.execute_reply":"2022-04-30T00:17:18.015533Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Number of training images of each species of whale and dolphin","metadata":{}},{"cell_type":"code","source":"fig,ax = plt.subplots(1,2,figsize=(10,6))\n\nwhales = train_df[train_df['class']=='whale']\ndolphins = train_df[train_df['class']!='whale']\n\nsns.countplot(y=\"species\", data=whales, order=whales.iloc[0:][\"species\"].value_counts().index, ax=ax[0], color = \"#0077b6\")\nax[0].set_title('Most frequent whales')\nax[0].set_ylabel(None)\n    \nsns.countplot(y=\"species\", data=dolphins,order=dolphins.iloc[0:][\"species\"].value_counts().index, ax=ax[1], color = \"#90e0ef\")\nax[1].set_title('Most frequent dolphins')\nax[1].set_ylabel(None)\n\nplt.tight_layout()\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-04-30T00:17:18.017846Z","iopub.execute_input":"2022-04-30T00:17:18.018115Z","iopub.status.idle":"2022-04-30T00:17:18.620954Z","shell.execute_reply.started":"2022-04-30T00:17:18.018077Z","shell.execute_reply":"2022-04-30T00:17:18.620201Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Number of training images of top 10 individuals","metadata":{}},{"cell_type":"code","source":"plt.figure(figsize=(12,4))\ntop_ten_ids = train_df.individual_id.value_counts().head(24)\ntop_ten_ids = pd.DataFrame({'individual_id':top_ten_ids.index, 'frequency':top_ten_ids.values})\n\nplt.bar(top_ten_ids['individual_id'],top_ten_ids['frequency'],width = 0.8,color='c',zorder=4)\nplt.xticks(rotation=90)\nplt.ylabel(\"frequency\")\nplt.xlabel(\"Individual Ids\")\nplt.title(\"Top 10 Individual Ids used by frequency\")\nplt.grid(visible = True, color ='grey',linestyle ='-', linewidth = 0.9,alpha = 0.2, zorder=0)\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-04-30T00:17:18.625397Z","iopub.execute_input":"2022-04-30T00:17:18.626072Z","iopub.status.idle":"2022-04-30T00:17:19.070338Z","shell.execute_reply.started":"2022-04-30T00:17:18.626031Z","shell.execute_reply":"2022-04-30T00:17:19.069637Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Plot the value count graph of each individual","metadata":{}},{"cell_type":"code","source":"train_df['individual_id'].value_counts().plot()\nplt.xticks(rotation=90)\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-04-30T00:17:19.071597Z","iopub.execute_input":"2022-04-30T00:17:19.072023Z","iopub.status.idle":"2022-04-30T00:17:19.289330Z","shell.execute_reply.started":"2022-04-30T00:17:19.071984Z","shell.execute_reply":"2022-04-30T00:17:19.288632Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Density estimation of each individuals","metadata":{}},{"cell_type":"code","source":"np.log(train_df['individual_id'].value_counts()).plot.kde()","metadata":{"execution":{"iopub.status.busy":"2022-04-30T00:17:19.290474Z","iopub.execute_input":"2022-04-30T00:17:19.291296Z","iopub.status.idle":"2022-04-30T00:17:19.789108Z","shell.execute_reply.started":"2022-04-30T00:17:19.291248Z","shell.execute_reply":"2022-04-30T00:17:19.788414Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Image count of individuals","metadata":{}},{"cell_type":"code","source":"train_df['count'] = train_df.groupby('individual_id',as_index=False)['individual_id'].transform(lambda x: x.count())\ntrain_df.head()","metadata":{"execution":{"iopub.status.busy":"2022-04-30T00:17:19.790269Z","iopub.execute_input":"2022-04-30T00:17:19.790509Z","iopub.status.idle":"2022-04-30T00:17:39.368712Z","shell.execute_reply.started":"2022-04-30T00:17:19.790471Z","shell.execute_reply":"2022-04-30T00:17:39.368015Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"observation-regarding-class-distribution\"></a>\n## Observation Regarding Class Distribution\nThere is a huge disbalance in the data. There are many classes with only one or several samples:\n\n1. Total Number of individuals are 15587\n2. 9258 individuals have just one image\n3. Single whale with most images have 400 of them\n4. Images dsitribution:\n  1. almost 40% comes from whales with 4 or less images.\n  1. almost 23% comes from whales with 5-20 images.\n  1. rest 37% comes from individual with >20 images.","metadata":{}},{"cell_type":"markdown","source":"<a id=\"image-resolutions\"></a>\n# Image Resolutions","metadata":{}},{"cell_type":"code","source":"'''widths, heights = [], []\n\nfor path in tqdm(train_df[\"path\"]):\n    width, height = Image.open(path).size\n    widths.append(width)\n    heights.append(height)\n    \ntrain_df[\"width\"] = widths\ntrain_df[\"height\"] = heights\ntrain_df[\"dimension\"] = train_df[\"width\"] * train_df[\"height\"]\ntrain_df_save = train_df.copy()'''","metadata":{"execution":{"iopub.status.busy":"2022-04-30T00:17:39.370000Z","iopub.execute_input":"2022-04-30T00:17:39.370403Z","iopub.status.idle":"2022-04-30T00:17:39.376918Z","shell.execute_reply.started":"2022-04-30T00:17:39.370364Z","shell.execute_reply":"2022-04-30T00:17:39.375985Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Lets see some small images","metadata":{}},{"cell_type":"code","source":"'''train_df.sort_values('width').head(84)'''","metadata":{"execution":{"iopub.status.busy":"2022-04-30T00:17:53.958293Z","iopub.execute_input":"2022-04-30T00:17:53.958630Z","iopub.status.idle":"2022-04-30T00:17:53.967188Z","shell.execute_reply.started":"2022-04-30T00:17:53.958590Z","shell.execute_reply":"2022-04-30T00:17:53.966030Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"preprocessing\"></a>\n# Preprocessing\n### Encoding Labels","metadata":{}},{"cell_type":"code","source":"from sklearn.preprocessing import LabelEncoder, OneHotEncoder\nfrom sklearn.compose import ColumnTransformer\n\nX = train_df.iloc[:, 3].values\ny = train_df.iloc[:, 2].values\n\nlabel_encoder = LabelEncoder()\ny = label_encoder.fit_transform(y)\nonehot_encoder = OneHotEncoder(sparse=False)\ny = y.reshape(len(y), 1)\ny = onehot_encoder.fit_transform(y)","metadata":{"execution":{"iopub.status.busy":"2022-04-30T00:18:02.837830Z","iopub.execute_input":"2022-04-30T00:18:02.838426Z","iopub.status.idle":"2022-04-30T00:18:03.222423Z","shell.execute_reply.started":"2022-04-30T00:18:02.838387Z","shell.execute_reply":"2022-04-30T00:18:03.221699Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"y.shape","metadata":{"execution":{"iopub.status.busy":"2022-04-30T00:18:03.495208Z","iopub.execute_input":"2022-04-30T00:18:03.495455Z","iopub.status.idle":"2022-04-30T00:18:03.501666Z","shell.execute_reply.started":"2022-04-30T00:18:03.495426Z","shell.execute_reply":"2022-04-30T00:18:03.500721Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Modeling","metadata":{}},{"cell_type":"code","source":"train_jpg_path = \"../input/happy-whale-and-dolphin/train_images\"\ntest_jpg_peth = \"../input/happy-whale-and-dolphin/test_images\"\ntrain_images_list = os.listdir('../input/happy-whale-and-dolphin/train_images')","metadata":{"execution":{"iopub.status.busy":"2022-04-30T00:18:05.231975Z","iopub.execute_input":"2022-04-30T00:18:05.232444Z","iopub.status.idle":"2022-04-30T00:18:05.255352Z","shell.execute_reply.started":"2022-04-30T00:18:05.232404Z","shell.execute_reply":"2022-04-30T00:18:05.254688Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def Loading_Images(data, m, dataset):\n    print(\"Loading images\")\n    X_train = np.zeros((m, 32, 32, 3))\n    count = 0\n    for fig in tqdm(data['image']):\n        img = image.load_img(\"../input/happy-whale-and-dolphin/\"+dataset+\"/\"+fig, target_size=(32, 32, 3))\n        x = image.img_to_array(img)\n        x = preprocess_input(x)\n        X_train[count] = x\n        count += 1\n    return X_train","metadata":{"execution":{"iopub.status.busy":"2022-04-30T00:18:05.427264Z","iopub.execute_input":"2022-04-30T00:18:05.427514Z","iopub.status.idle":"2022-04-30T00:18:05.433340Z","shell.execute_reply.started":"2022-04-30T00:18:05.427487Z","shell.execute_reply":"2022-04-30T00:18:05.432686Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def prepare_labels(y):\n    values = np.array(y)\n    label_encoder = LabelEncoder()\n    integer_encoded = label_encoder.fit_transform(values)\n    onehot_encoder = OneHotEncoder(sparse=False)\n    integer_encoded = integer_encoded.reshape(len(integer_encoded), 1)\n    onehot_encoded = onehot_encoder.fit_transform(integer_encoded)\n    y = onehot_encoded\n    return y, label_encoder","metadata":{"execution":{"iopub.status.busy":"2022-04-30T00:18:06.329927Z","iopub.execute_input":"2022-04-30T00:18:06.330433Z","iopub.status.idle":"2022-04-30T00:18:06.335315Z","shell.execute_reply.started":"2022-04-30T00:18:06.330392Z","shell.execute_reply":"2022-04-30T00:18:06.334664Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"X = Loading_Images(train_df, train_df.shape[0], \"train_images\")\nX /= 255","metadata":{"execution":{"iopub.status.busy":"2022-04-30T00:33:53.222464Z","iopub.execute_input":"2022-04-30T00:33:53.222741Z","iopub.status.idle":"2022-04-30T01:41:14.887655Z","shell.execute_reply.started":"2022-04-30T00:33:53.222709Z","shell.execute_reply":"2022-04-30T01:41:14.886791Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import gc\ny, label_encoder = prepare_labels(train_df['individual_id'])\nprint(X.shape)\nprint(y.shape)\ngc.collect()","metadata":{"execution":{"iopub.status.busy":"2022-04-30T01:41:14.891298Z","iopub.execute_input":"2022-04-30T01:41:14.892363Z","iopub.status.idle":"2022-04-30T01:41:15.495281Z","shell.execute_reply.started":"2022-04-30T01:41:14.892319Z","shell.execute_reply":"2022-04-30T01:41:15.494552Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"y.shape","metadata":{"execution":{"iopub.status.busy":"2022-04-30T01:41:15.496626Z","iopub.execute_input":"2022-04-30T01:41:15.497071Z","iopub.status.idle":"2022-04-30T01:41:15.502231Z","shell.execute_reply.started":"2022-04-30T01:41:15.497031Z","shell.execute_reply":"2022-04-30T01:41:15.501485Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from tensorflow.keras.applications import EfficientNetB0\nfrom tensorflow.keras.layers import GlobalAveragePooling2D, Dropout, Dense\nfrom tensorflow.keras.preprocessing.image import ImageDataGenerator\nimport tensorflow as tf\nfrom keras.models import Model\n\n\nbase_model = EfficientNetB0(input_shape=(32,32,3), weights=None, include_top=False)\n\nlayer = base_model.output\n#layer = GlobalAveragePooling2D()(layer)#extra\n#layer = Dropout(0.5)(layer)#extra\nlayer = Dense(1024, activation='relu')(layer)\n#layer = Dense(512, activation='relu')(layer)#extra\nlayer = Flatten()(layer)\npredictions = Dense(y.shape[1], activation='softmax')(layer)\nmodel = Model(inputs=base_model.input, outputs=predictions)\n\nmodel.compile(loss='categorical_crossentropy', optimizer=\"adam\", metrics=['accuracy'])\nmodel.summary()","metadata":{"execution":{"iopub.status.busy":"2022-04-30T01:41:15.504225Z","iopub.execute_input":"2022-04-30T01:41:15.504624Z","iopub.status.idle":"2022-04-30T01:41:17.129336Z","shell.execute_reply.started":"2022-04-30T01:41:15.504586Z","shell.execute_reply":"2022-04-30T01:41:17.128636Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_datagen = ImageDataGenerator(horizontal_flip=True,\n                                   vertical_flip=True,\n                                   validation_split=0.20,\n                                   )\n\n#train_datagen.fit(X)","metadata":{"execution":{"iopub.status.busy":"2022-04-30T01:41:17.130441Z","iopub.execute_input":"2022-04-30T01:41:17.130771Z","iopub.status.idle":"2022-04-30T01:41:17.135195Z","shell.execute_reply.started":"2022-04-30T01:41:17.130724Z","shell.execute_reply":"2022-04-30T01:41:17.134223Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#history = model.fit(train_datagen.flow(X,y,batch_size=128,subset='training'),validation_data=train_datagen.flow(X,y,batch_size=128,subset='validation'),epochs=180)\nhistory = model.fit(X, y, epochs = 200, batch_size=128, verbose=1)","metadata":{"execution":{"iopub.status.busy":"2022-04-30T01:42:26.316306Z","iopub.execute_input":"2022-04-30T01:42:26.316562Z","iopub.status.idle":"2022-04-30T02:56:23.322272Z","shell.execute_reply.started":"2022-04-30T01:42:26.316532Z","shell.execute_reply":"2022-04-30T02:56:23.321548Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def cnn_model():\n    model = Sequential()\n    model.add(Conv2D(32, (6, 6), strides = (1, 1), input_shape = (32, 32, 3)))\n    model.add(BatchNormalization(axis = 3))\n    model.add(Activation('relu'))\n    model.add(MaxPooling2D((2, 2)))\n      \n    model.add(Conv2D(64, (3, 3), strides = (1,1)))\n    model.add(Activation('relu'))\n    model.add(AveragePooling2D((3, 3)))\n\n    model.add(Flatten())\n    model.add(Dense(512, activation=\"relu\"))\n    model.add(Dropout(0.85))\n\n    model.add(Dense(y.shape[1], activation='softmax'))\n\n    model.compile(loss='categorical_crossentropy', optimizer=\"adam\", metrics=['accuracy'])\n    \n    return(model)","metadata":{"execution":{"iopub.status.busy":"2022-04-30T03:09:08.138248Z","iopub.execute_input":"2022-04-30T03:09:08.138529Z","iopub.status.idle":"2022-04-30T03:09:08.146767Z","shell.execute_reply.started":"2022-04-30T03:09:08.138487Z","shell.execute_reply":"2022-04-30T03:09:08.145714Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"Cnn_model = cnn_model()","metadata":{"execution":{"iopub.status.busy":"2022-04-30T03:09:08.148338Z","iopub.execute_input":"2022-04-30T03:09:08.148785Z","iopub.status.idle":"2022-04-30T03:09:09.544547Z","shell.execute_reply.started":"2022-04-30T03:09:08.148747Z","shell.execute_reply":"2022-04-30T03:09:09.543527Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.figure(figsize=(15,5))\nplt.plot(history.history['accuracy'])\nplt.title('Model accuracy')\nplt.ylabel('Accuracy')\nplt.xlabel('Epoch')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-04-30T03:08:25.074090Z","iopub.execute_input":"2022-04-30T03:08:25.074793Z","iopub.status.idle":"2022-04-30T03:08:25.090490Z","shell.execute_reply.started":"2022-04-30T03:08:25.074746Z","shell.execute_reply":"2022-04-30T03:08:25.089409Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.figure(figsize=(15,5))\nplt.plot(history.history['loss'])\nplt.title('Model loss')\nplt.ylabel('loss')\nplt.xlabel('Epoch')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-04-30T02:59:12.546401Z","iopub.status.idle":"2022-04-30T02:59:12.547176Z","shell.execute_reply.started":"2022-04-30T02:59:12.546939Z","shell.execute_reply":"2022-04-30T02:59:12.546968Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test = os.listdir(\"../input/happy-whale-and-dolphin/test_images\")\nprint(len(test))","metadata":{"execution":{"iopub.status.busy":"2022-04-30T02:59:12.548230Z","iopub.status.idle":"2022-04-30T02:59:12.549965Z","shell.execute_reply.started":"2022-04-30T02:59:12.549717Z","shell.execute_reply":"2022-04-30T02:59:12.549745Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"col = ['image']\ntest_df = pd.DataFrame(test, columns=col)\ntest_df['predictions'] = ''","metadata":{"execution":{"iopub.status.busy":"2022-04-30T02:59:12.551171Z","iopub.status.idle":"2022-04-30T02:59:12.551582Z","shell.execute_reply.started":"2022-04-30T02:59:12.551365Z","shell.execute_reply":"2022-04-30T02:59:12.551387Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"batch_size=5000\nbatch_start = 0\nbatch_end = batch_size\nL = len(test_df)\n\nwhile batch_start < L:\n    limit = min(batch_end, L)\n    test_df_batch = test_df.iloc[batch_start:limit]\n    print(type(test_df_batch))\n    X = Loading_Images(test_df_batch, test_df_batch.shape[0], \"test_images\")\n    X /= 255\n    predictions = model.predict(np.array(X), verbose=1)\n    for i, pred in enumerate(predictions):\n        p=pred.argsort()[-5:][::-1]\n        idx=-1\n        s=''\n        s1=''\n        s2=''\n        for x in p:\n            idx=idx+1\n            if pred[x]>0.5:\n                s1 = s1 + ' ' +  label_encoder.inverse_transform(p)[idx]\n            else:\n                s2 = s2 + ' ' + label_encoder.inverse_transform(p)[idx]\n        s= s1 + ' new_individual' + s2\n        s = s.strip(' ')\n        test_df.loc[ batch_start + i, 'predictions'] = s\n    batch_start += batch_size   \n    batch_end += batch_size\n    del X\n    del test_df_batch\n    del predictions\n    gc.collect()","metadata":{"execution":{"iopub.status.busy":"2022-04-30T02:59:12.552894Z","iopub.status.idle":"2022-04-30T02:59:12.553633Z","shell.execute_reply.started":"2022-04-30T02:59:12.553388Z","shell.execute_reply":"2022-04-30T02:59:12.553412Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_df.to_csv('submission.csv',index=False)\ntest_df.head()","metadata":{"execution":{"iopub.status.busy":"2022-04-30T02:59:12.554951Z","iopub.status.idle":"2022-04-30T02:59:12.555366Z","shell.execute_reply.started":"2022-04-30T02:59:12.555130Z","shell.execute_reply":"2022-04-30T02:59:12.555153Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_df.to_csv('submission_whale_and_dolphin.csv', index = False)","metadata":{"execution":{"iopub.status.busy":"2022-04-30T02:59:12.538518Z","iopub.status.idle":"2022-04-30T02:59:12.539169Z","shell.execute_reply.started":"2022-04-30T02:59:12.538923Z","shell.execute_reply":"2022-04-30T02:59:12.538948Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"","metadata":{}},{"cell_type":"markdown","source":"# References\nI have used these awesome kernels for whole EDA. Do check them out if you have time.","metadata":{}},{"cell_type":"code","source":"##https://www.kaggle.com/andradaolteanu/whales-dolphins-effnet-embedding-cos-distance\n##https://www.kaggle.com/lextoumbourou/happy-whale-dolphin-q-a-style-eda\n##https://www.kaggle.com/andradaolteanu/whales-dolphins-effnet-embedding-cos-distance\n##https://www.kaggle.com/rednivrug/eda-for-whale-with-bounding-boxes/notebook\n##https://www.kaggle.com/andradaolteanu/whales-dolphins-effnet-embedding-cos-distance\n##https://www.kaggle.com/pestipeti/explanation-of-map5-scoring-metric","metadata":{"execution":{"iopub.status.busy":"2022-04-30T00:17:39.726871Z","iopub.status.idle":"2022-04-30T00:17:39.727418Z","shell.execute_reply.started":"2022-04-30T00:17:39.727181Z","shell.execute_reply":"2022-04-30T00:17:39.727204Z"},"trusted":true},"execution_count":null,"outputs":[]}]}