{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# HAPPY WHALES AND DOLPHINS IDENTIFICATION - CNN Model","metadata":{}},{"cell_type":"markdown","source":"A CNN Model","metadata":{}},{"cell_type":"markdown","source":"Competition Link: https://www.kaggle.com/competitions/happy-whale-and-dolphin\n\nData Source: https://www.kaggle.com/competitions/happy-whale-and-dolphin/data\n\n\n#### Files:\n\n* train_images/ - a folder containing the training images\n* train.csv - provides the species and the individual_id for each of the training images\n* test_images/ - a folder containing the test images; for each image, the task is to predict the individual_id; no species information is given for the test data; there are individuals in the test data that are not observed in the training data, which should be predicted as new_individual.\n* sample_submission.csv - a sample submission file in the correct format\n\n#### This notebook has the follow steps:\n\n* Brief Description of the Problem and Data\n* Exploratory Data Analysis (EDA) and Preprocessing\n* Model Construction and Architecture\n* Results and Analysis\n* Conclusion","metadata":{}},{"cell_type":"code","source":" # Import libraries:\n\nimport numpy as np\nimport pandas as pd\nimport os\nimport random\nimport shutil\nimport glob\nfrom sklearn.utils import shuffle\n\n# for image:\nimport seaborn as sns\nimport matplotlib.pyplot as plt\n%matplotlib inline\nimport matplotlib.image as mpimg\n\nfrom sklearn.preprocessing import LabelEncoder\nfrom sklearn.preprocessing import OneHotEncoder\n\n# for model:\nimport tensorflow as tf\nfrom tensorflow import keras\nfrom tensorflow.keras.models import Sequential\nfrom tensorflow.keras.layers import Dense, Flatten, Activation, Dropout, BatchNormalization, LeakyReLU\nfrom tensorflow.keras.layers import Conv2D, AveragePooling2D, MaxPooling2D\nfrom tensorflow.keras.optimizers import Adam\nfrom tensorflow.keras.preprocessing.image import ImageDataGenerator","metadata":{"execution":{"iopub.status.busy":"2023-03-07T22:11:42.243409Z","iopub.execute_input":"2023-03-07T22:11:42.244071Z","iopub.status.idle":"2023-03-07T22:11:51.858524Z","shell.execute_reply.started":"2023-03-07T22:11:42.244025Z","shell.execute_reply":"2023-03-07T22:11:51.857155Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Check tensorflow verison:\n\ntf.__version__","metadata":{"execution":{"iopub.status.busy":"2023-03-07T22:11:51.860029Z","iopub.execute_input":"2023-03-07T22:11:51.860766Z","iopub.status.idle":"2023-03-07T22:11:51.869297Z","shell.execute_reply.started":"2023-03-07T22:11:51.860727Z","shell.execute_reply":"2023-03-07T22:11:51.868171Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Brief Description of the Problem and Data:\n\nWhales and dolphins are a key component for the marine ecosystem so being able to identify species could be helpful for other research projects. \n\nIn this project, we are given 51033 images which are connected to the training dataframe, in which we are given the columns 'image id' as jpg, 'species', and 'individual_id'. We also are given testing images.\n\nWe also note that we do not have bounding boxes on the images and all the images are different shapes and resolutions.\n\nThe main task was to predict individual whales and dolphins from images of their fins. Whales and dolphins in this dataset can be identified by shapes, features and markings of dorsal fins, backs, heads and flanks.\n","metadata":{}},{"cell_type":"code","source":"# Get global path:\n\nprint(os.listdir('../input'))","metadata":{"execution":{"iopub.status.busy":"2023-03-07T22:11:51.870561Z","iopub.execute_input":"2023-03-07T22:11:51.871640Z","iopub.status.idle":"2023-03-07T22:11:51.881929Z","shell.execute_reply.started":"2023-03-07T22:11:51.871583Z","shell.execute_reply":"2023-03-07T22:11:51.880687Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Get global path:\n\nprint(os.listdir('../input/happy-whale-and-dolphin'))","metadata":{"execution":{"iopub.status.busy":"2023-03-07T22:11:51.886488Z","iopub.execute_input":"2023-03-07T22:11:51.887473Z","iopub.status.idle":"2023-03-07T22:11:51.895946Z","shell.execute_reply.started":"2023-03-07T22:11:51.887414Z","shell.execute_reply":"2023-03-07T22:11:51.894625Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Set paths:\n\ntrain = '../input/happy-whale-and-dolphin/train_images'\ntest = '../input/happy-whale-and-dolphin/test_images'","metadata":{"execution":{"iopub.status.busy":"2023-03-07T22:11:51.897263Z","iopub.execute_input":"2023-03-07T22:11:51.897576Z","iopub.status.idle":"2023-03-07T22:11:51.905222Z","shell.execute_reply.started":"2023-03-07T22:11:51.897546Z","shell.execute_reply":"2023-03-07T22:11:51.904267Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Dimensions:\n\nprint(len(os.listdir(train)))\nprint(len(os.listdir(test)))","metadata":{"execution":{"iopub.status.busy":"2023-03-07T22:11:51.906460Z","iopub.execute_input":"2023-03-07T22:11:51.906791Z","iopub.status.idle":"2023-03-07T22:11:56.040297Z","shell.execute_reply.started":"2023-03-07T22:11:51.906759Z","shell.execute_reply":"2023-03-07T22:11:56.039456Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Get image paths:\n\nprint(os.listdir(train)[:5])\nprint(os.listdir(test)[:5])","metadata":{"execution":{"iopub.status.busy":"2023-03-07T22:11:56.041572Z","iopub.execute_input":"2023-03-07T22:11:56.042114Z","iopub.status.idle":"2023-03-07T22:11:56.076077Z","shell.execute_reply.started":"2023-03-07T22:11:56.042079Z","shell.execute_reply":"2023-03-07T22:11:56.074649Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Set image paths:\n\ntrain_jpg = tf.io.gfile.glob(train+'/*.jpg')\ntest_jpg = tf.io.gfile.glob(test+'/*.jpg')","metadata":{"execution":{"iopub.status.busy":"2023-03-07T22:11:56.077545Z","iopub.execute_input":"2023-03-07T22:11:56.078040Z","iopub.status.idle":"2023-03-07T22:12:50.527147Z","shell.execute_reply.started":"2023-03-07T22:11:56.077988Z","shell.execute_reply":"2023-03-07T22:12:50.526187Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# View train dataset:\n\ntrain_data = pd.read_csv('../input/happy-whale-and-dolphin/train.csv', sep = ',')\ntrain_data.head()","metadata":{"execution":{"iopub.status.busy":"2023-03-07T22:12:50.528688Z","iopub.execute_input":"2023-03-07T22:12:50.529398Z","iopub.status.idle":"2023-03-07T22:12:50.665741Z","shell.execute_reply.started":"2023-03-07T22:12:50.529352Z","shell.execute_reply":"2023-03-07T22:12:50.664410Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# View train dataset:\n\ntrain_data.sample(5)","metadata":{"execution":{"iopub.status.busy":"2023-03-07T22:12:50.667347Z","iopub.execute_input":"2023-03-07T22:12:50.669883Z","iopub.status.idle":"2023-03-07T22:12:50.686115Z","shell.execute_reply.started":"2023-03-07T22:12:50.669844Z","shell.execute_reply":"2023-03-07T22:12:50.685209Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Size:\n\ntrain_data.shape","metadata":{"execution":{"iopub.status.busy":"2023-03-07T22:12:50.687425Z","iopub.execute_input":"2023-03-07T22:12:50.687839Z","iopub.status.idle":"2023-03-07T22:12:50.694727Z","shell.execute_reply.started":"2023-03-07T22:12:50.687806Z","shell.execute_reply":"2023-03-07T22:12:50.693562Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Training data information and data types:\n\ntrain_data.info()","metadata":{"execution":{"iopub.status.busy":"2023-03-07T22:12:50.696408Z","iopub.execute_input":"2023-03-07T22:12:50.696914Z","iopub.status.idle":"2023-03-07T22:12:50.729571Z","shell.execute_reply.started":"2023-03-07T22:12:50.696866Z","shell.execute_reply":"2023-03-07T22:12:50.728440Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Data Description:\n\ntrain_data.describe(include = 'all')","metadata":{"execution":{"iopub.status.busy":"2023-03-07T22:12:50.735998Z","iopub.execute_input":"2023-03-07T22:12:50.736427Z","iopub.status.idle":"2023-03-07T22:12:50.793802Z","shell.execute_reply.started":"2023-03-07T22:12:50.736394Z","shell.execute_reply":"2023-03-07T22:12:50.792526Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Number of missing data:\n\ntrain_data.isnull().sum()","metadata":{"execution":{"iopub.status.busy":"2023-03-07T22:12:50.795014Z","iopub.execute_input":"2023-03-07T22:12:50.795299Z","iopub.status.idle":"2023-03-07T22:12:50.812297Z","shell.execute_reply.started":"2023-03-07T22:12:50.795271Z","shell.execute_reply":"2023-03-07T22:12:50.811172Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Value counts of species variable:\n\ntrain_data.species.value_counts()","metadata":{"execution":{"iopub.status.busy":"2023-03-07T22:12:50.813899Z","iopub.execute_input":"2023-03-07T22:12:50.814278Z","iopub.status.idle":"2023-03-07T22:12:50.826180Z","shell.execute_reply.started":"2023-03-07T22:12:50.814245Z","shell.execute_reply":"2023-03-07T22:12:50.824849Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Check duplication:\n\nsum(train_data.individual_id.duplicated())","metadata":{"execution":{"iopub.status.busy":"2023-03-07T22:12:50.827863Z","iopub.execute_input":"2023-03-07T22:12:50.828195Z","iopub.status.idle":"2023-03-07T22:12:50.847204Z","shell.execute_reply.started":"2023-03-07T22:12:50.828163Z","shell.execute_reply":"2023-03-07T22:12:50.845953Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Exploratory Data Analysis (EDA) and Preprocessing:\n\nSummary of preprocessing steps we will perform:\n\n* Loading the image\n* Fix misspellings\n* Properly group whales and dolphins\n* Normalize images (multiply by 1/255)\n* Crop images to 64x64 pixel and 3 channels (64, 64, 3)\n* Split into training-validation sets (80-20 split)","metadata":{}},{"cell_type":"code","source":"# Loading an image files by its path:\n\ndef Load_Image(path):\n    image_path = tf.io.read_file(path)\n    image_path = tf.image.decode_image(image_path, channels = 3)\n    image_path = tf.image.convert_image_dtype(image_path, tf.float32)\n    return image_path","metadata":{"execution":{"iopub.status.busy":"2023-03-07T22:12:50.848184Z","iopub.execute_input":"2023-03-07T22:12:50.848527Z","iopub.status.idle":"2023-03-07T22:12:50.856214Z","shell.execute_reply.started":"2023-03-07T22:12:50.848477Z","shell.execute_reply":"2023-03-07T22:12:50.855068Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Visualization:\n# Plot of unique species and their percentage:\n\n\nplt.figure(figsize = [20,6])\n\ncolor = sns.color_palette()[0]\n\nsns.countplot(data = train_data, x = 'species', color = color);\nplt.title('Simple Plot');\nplt.ylabel('count');\nplt.xlabel('species');\n\nvalue_sum = train_data['species'].value_counts().sum()\nvalue = train_data['species'].value_counts()\n\nlocs, labels = plt.xticks(rotation = 90) \n\nfor loc, label in zip(locs, labels):\n\n    count = value[label.get_text()]\n    text = '{:0.1f}%'.format(100 * count/value_sum)\n\n    plt.text(loc, count+3, text, ha = 'center', color = 'black');","metadata":{"execution":{"iopub.status.busy":"2023-03-07T22:12:50.858032Z","iopub.execute_input":"2023-03-07T22:12:50.858353Z","iopub.status.idle":"2023-03-07T22:12:51.920785Z","shell.execute_reply.started":"2023-03-07T22:12:50.858317Z","shell.execute_reply":"2023-03-07T22:12:51.919606Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Fix mis-spellings from species variable:\n\ntrain_data['species'] = train_data['species'].replace({'kiler_whale': 'killer_whale', \n                               'globis': 'pilot_whale', \n                               'beluga': 'beluga_whale',\n                               'bottlenose_dolpin': 'bottlenose_dolphin',\n                               'short_finned_pilot_whale': 'pilot_whale',\n                               'long_finned_pilot_whale': 'pilot_whale'})","metadata":{"execution":{"iopub.status.busy":"2023-03-07T22:12:51.922164Z","iopub.execute_input":"2023-03-07T22:12:51.922595Z","iopub.status.idle":"2023-03-07T22:12:51.949612Z","shell.execute_reply.started":"2023-03-07T22:12:51.922547Z","shell.execute_reply":"2023-03-07T22:12:51.948487Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Declare dolphin and whale variables for further analysis:\n\ndolphin = ['bottlenose_dolphin','common_dolphin','dusky_dolphin', 'spinner_dolphin', 'spotted_dolphin', 'commersons_dolphin', \n           'white_sided_dolphin', 'rough_toothed_dolphin', 'pantropic_spotted_dolphin', 'frasiers_dolphin']\n\n\nwhale = ['melon_headed_whale', 'humpback_whale', 'false_killer_whale', 'belug_whale', 'minke_whale', 'fin_whale', 'blue_whale', 'gray_whale',\n         'southern_right_whale', 'killer_whale', 'pilot_whale', 'sei_whale', 'cuviers_beaked_whale', 'brydes_whale', 'pygmy_killer_whale']\n\n\n# Add to train dataset:\ntrain_data['family'] = 'dolphin'\n\nfor ele in range(len(train_data)):\n    if train_data.species[ele] in whale:\n        train_data.family[ele] = 'whale'\n        \n        \ntrain_data.sample(5)","metadata":{"execution":{"iopub.status.busy":"2023-03-07T22:12:51.951334Z","iopub.execute_input":"2023-03-07T22:12:51.952033Z","iopub.status.idle":"2023-03-07T22:13:04.916764Z","shell.execute_reply.started":"2023-03-07T22:12:51.951989Z","shell.execute_reply":"2023-03-07T22:13:04.915492Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Visualization:\n# Plot of family of species and their percentage:\n\nplt.figure(figsize = [10, 5])\nlabels = train_data.family.value_counts()\n\nplt.subplot(1,2,1)\nlabels.plot(kind = 'pie', autopct = '%1.2f%%', shadow = True, startangle = 180)\nplt.title('Simple Pie Plot',fontsize = 15)\nplt.ylabel('Family',fontsize = 10);\n\nplt.subplot(1,2,2)\nsns.countplot(data = train_data, x = 'family', color = color);\nplt.title('Simple Plot');\nplt.ylabel('count');\nplt.xlabel('Family');","metadata":{"execution":{"iopub.status.busy":"2023-03-07T22:13:04.918338Z","iopub.execute_input":"2023-03-07T22:13:04.919253Z","iopub.status.idle":"2023-03-07T22:13:05.308019Z","shell.execute_reply.started":"2023-03-07T22:13:04.919218Z","shell.execute_reply":"2023-03-07T22:13:05.306884Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Visualization:\n# Plot of unique species and their percentage with hue = family:\n\nplt.figure(figsize = [20,6])\n\nsns.countplot(data = train_data, x = 'species', hue = 'family', order = train_data['species'].value_counts().index);\nplt.title('Simple Plot');\nplt.ylabel('count');\nplt.xlabel('species');\n\nvalue_sum = train_data['species'].value_counts().sum()\nvalue = train_data['species'].value_counts()\n\nlocs, labels = plt.xticks(rotation = 90) \n\nfor loc, label in zip(locs, labels):\n\n    count = value[label.get_text()]\n    text = '{:0.1f}%'.format(100 * count/value_sum)\n\n    plt.text(loc, count+1, text, ha = 'center', color = 'black');","metadata":{"execution":{"iopub.status.busy":"2023-03-07T22:13:05.309705Z","iopub.execute_input":"2023-03-07T22:13:05.310680Z","iopub.status.idle":"2023-03-07T22:13:06.199071Z","shell.execute_reply.started":"2023-03-07T22:13:05.310642Z","shell.execute_reply":"2023-03-07T22:13:06.197828Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Visualization:\n# Plotting an example from train image:\n\nfrom keras.preprocessing import image\nfrom keras.utils import load_img, img_to_array\n\n\nplt.figure(figsize = (20, 20))\n\nplt.subplot(1, 4, 1)\nimage = load_img('../input/happy-whale-and-dolphin/train_images/80b5373b87942b.jpg')\nplt.imshow(image);","metadata":{"execution":{"iopub.status.busy":"2023-03-07T22:13:06.200850Z","iopub.execute_input":"2023-03-07T22:13:06.201184Z","iopub.status.idle":"2023-03-07T22:13:08.406480Z","shell.execute_reply.started":"2023-03-07T22:13:06.201151Z","shell.execute_reply":"2023-03-07T22:13:08.405117Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Visualization:\n# Plotting some examples from Train image path:\n\nfig, ax = plt.subplots(4, 5, figsize = (20, 20))\n\njpg = random.sample(train_jpg, 20)\n\nfor idx, name in enumerate(jpg):\n    img = Load_Image(name)\n    ax[idx//5, idx%5].imshow(img)\n    ax[idx//5, idx%5].set_title('Train image')","metadata":{"execution":{"iopub.status.busy":"2023-03-07T22:13:08.407774Z","iopub.execute_input":"2023-03-07T22:13:08.408098Z","iopub.status.idle":"2023-03-07T22:13:43.066051Z","shell.execute_reply.started":"2023-03-07T22:13:08.408067Z","shell.execute_reply":"2023-03-07T22:13:43.064280Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Visualization:\n# Plotting some examples from Test image path:\n\nfig, ax = plt.subplots(4, 5, figsize = (20, 20))\n\njpg = random.sample(test_jpg, 20)\n\nfor idx, name in enumerate(jpg):\n    img = Load_Image(name)\n    ax[idx//5, idx%5].imshow(img)\n    ax[idx//5, idx%5].set_title('Test image')","metadata":{"execution":{"iopub.status.busy":"2023-03-07T22:13:43.067773Z","iopub.execute_input":"2023-03-07T22:13:43.068862Z","iopub.status.idle":"2023-03-07T22:14:14.021977Z","shell.execute_reply.started":"2023-03-07T22:13:43.068809Z","shell.execute_reply":"2023-03-07T22:14:14.020748Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Visualization:\n# Visualization of Unique Specie:\n\ntrain_data['path'] = '../input/happy-whale-and-dolphin/train_images/' + train_data['image']\n\ndef species_plot(data, variable):\n    plt.figure(figsize = (12, 12))\n    df = data[data['species'] == variable].reset_index(drop = True)\n    plt.suptitle(variable)\n    \n    for idx, ele in enumerate(np.random.choice(df['path'], 16)):\n        plt.subplot(4, 4, idx+1)\n        image_path = ele\n        img = Load_Image(image_path)\n\n        plt.imshow(img)\n        \n    plt.tight_layout()\n    plt.show()\n    \n\nfor var in train_data['species'].unique()[:5]:\n    species_plot(train_data, var)\n    ","metadata":{"execution":{"iopub.status.busy":"2023-03-07T22:14:14.023640Z","iopub.execute_input":"2023-03-07T22:14:14.024017Z","iopub.status.idle":"2023-03-07T22:15:33.130250Z","shell.execute_reply.started":"2023-03-07T22:14:14.023981Z","shell.execute_reply":"2023-03-07T22:15:33.129073Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Model Construction and Architecture:","metadata":{}},{"cell_type":"markdown","source":"During model training, we shuffle the data, normalize the images, and crop them to 64x64 pixels. These steps can help make the model more robust and learn better because the images are all standardized and shown in more positions. We will be using Tensorflow layers to construct our models.\n\nFor this project we will use one very simple model. The model will have five layers after its \"base\" layer. We call it a \"base\" layer so that we may more easily compare the model later on. In addition, model's input will be with a shape of (64, 64, 3). And it has following layers:\n\n* Average Pooling\n* Dropout layer with 0.1 dropout\n* BatchNormalization\n* Output layer using Dense with unique categories and activation function as 'softmax'\n* Optimizer Adam","metadata":{}},{"cell_type":"code","source":"# Set globals:\n\nrandom_state = 42\nbatch_size = 256\nepochs = 3\nseed = 42\ntarget_size = (64, 64)\ninput_shape = (64, 64, 3)","metadata":{"execution":{"iopub.status.busy":"2023-03-07T22:15:33.131588Z","iopub.execute_input":"2023-03-07T22:15:33.132043Z","iopub.status.idle":"2023-03-07T22:15:33.137033Z","shell.execute_reply.started":"2023-03-07T22:15:33.132010Z","shell.execute_reply":"2023-03-07T22:15:33.135977Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Shuffle the dataset:\n\ntrain_data = shuffle(train_data, random_state = random_state)","metadata":{"execution":{"iopub.status.busy":"2023-03-07T22:15:33.138349Z","iopub.execute_input":"2023-03-07T22:15:33.139263Z","iopub.status.idle":"2023-03-07T22:15:33.165073Z","shell.execute_reply.started":"2023-03-07T22:15:33.139228Z","shell.execute_reply":"2023-03-07T22:15:33.163910Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Set up image generator to split into 80/20 train-validation groups:\n\ndata_norm = ImageDataGenerator(rescale = 1.0/255, validation_split = 0.20)","metadata":{"execution":{"iopub.status.busy":"2023-03-07T22:15:33.166444Z","iopub.execute_input":"2023-03-07T22:15:33.166831Z","iopub.status.idle":"2023-03-07T22:15:33.172277Z","shell.execute_reply.started":"2023-03-07T22:15:33.166794Z","shell.execute_reply":"2023-03-07T22:15:33.171094Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Eet up training batching:\n\ngen_train = data_norm.flow_from_dataframe(train_data,\n                                          directory = train,\n                                          x_col = 'image',\n                                          y_col = 'species',\n                                          subset = 'training',\n                                          batch_size = batch_size,\n                                          class_mode = 'categorical',\n                                          seed = seed,\n                                          target_size = target_size)","metadata":{"execution":{"iopub.status.busy":"2023-03-07T22:15:33.174278Z","iopub.execute_input":"2023-03-07T22:15:33.174741Z","iopub.status.idle":"2023-03-07T22:15:58.020360Z","shell.execute_reply.started":"2023-03-07T22:15:33.174694Z","shell.execute_reply":"2023-03-07T22:15:58.019201Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Set up testing/validation batching:\n\ngen_valid = data_norm.flow_from_dataframe(train_data,\n                                          directory = train,\n                                          x_col = 'image',\n                                          y_col = 'species',\n                                          subset = 'validation',\n                                          batch_size = batch_size,\n                                          class_mode = 'categorical',\n                                          seed = seed,\n                                          target_size = target_size)","metadata":{"execution":{"iopub.status.busy":"2023-03-07T22:15:58.022069Z","iopub.execute_input":"2023-03-07T22:15:58.022525Z","iopub.status.idle":"2023-03-07T22:16:23.802987Z","shell.execute_reply.started":"2023-03-07T22:15:58.022478Z","shell.execute_reply":"2023-03-07T22:16:23.801698Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Build and Train Simple Model:\n\nmod = Sequential()\n\n# set up base model (simple base):\nmod.add(Conv2D(filters = 32, kernel_size = (5, 5), strides = (1, 1), input_shape = input_shape, padding ='valid'))\nmod.add(BatchNormalization())\nmod.add(Activation(LeakyReLU()))\n\nmod.add(Conv2D(filters = 32, kernel_size = (5, 5), strides = (1, 1), input_shape = input_shape, padding ='valid'))\nmod.add(BatchNormalization())\nmod.add(Activation(LeakyReLU()))\nmod.add(MaxPooling2D(pool_size = (2, 2)))\nmod.add(Dropout(0.1))\n\nmod.add(Conv2D(filters = 64, kernel_size = (5, 5), strides = (1, 1), input_shape = input_shape, padding ='valid'))\nmod.add(Activation(LeakyReLU()))\nmod.add(BatchNormalization())\n\nmod.add(Conv2D(filters = 128, kernel_size = (5, 5), strides = (1, 1), input_shape = input_shape, padding ='valid'))\nmod.add(BatchNormalization())\nmod.add(Activation(LeakyReLU()))\nmod.add(AveragePooling2D(pool_size = (2, 2)))\n\nmod.add(Conv2D(filters = 128, kernel_size = (5, 5), strides = (1, 1), input_shape = input_shape, padding ='valid'))\nmod.add(BatchNormalization())\nmod.add(Activation(LeakyReLU()))\nmod.add(AveragePooling2D(pool_size = (2, 2)))\nmod.add(Dropout(0.1))\n\nmod.add(Flatten())\n\n# set dense with activation as softmax:\nmod.add(Dense(train_data.species.nunique(), activation = 'softmax'))\n\n# set optimizer with small rate:\nopt = Adam(learning_rate = 0.0001)\n\n#set up loss function:\nlosses = tf.keras.losses.CategoricalCrossentropy() \n\n# compile model:\nmod.compile(loss = 'categorical_crossentropy', metrics = ['accuracy'], optimizer = opt)\n\n# view model summary:\nmod.summary()","metadata":{"execution":{"iopub.status.busy":"2023-03-07T22:16:23.804817Z","iopub.execute_input":"2023-03-07T22:16:23.805348Z","iopub.status.idle":"2023-03-07T22:16:24.179673Z","shell.execute_reply.started":"2023-03-07T22:16:23.805300Z","shell.execute_reply":"2023-03-07T22:16:24.178313Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Train model:\n\nfit = mod.fit(gen_train, epochs = epochs, validation_data = gen_valid)\nfit","metadata":{"execution":{"iopub.status.busy":"2023-03-07T22:16:24.181815Z","iopub.execute_input":"2023-03-07T22:16:24.182288Z","iopub.status.idle":"2023-03-08T02:45:37.188970Z","shell.execute_reply.started":"2023-03-07T22:16:24.182240Z","shell.execute_reply":"2023-03-08T02:45:37.183866Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Check keys before plotting:\n\nfit.history","metadata":{"execution":{"iopub.status.busy":"2023-03-08T02:45:37.196453Z","iopub.execute_input":"2023-03-08T02:45:37.197256Z","iopub.status.idle":"2023-03-08T02:45:37.210143Z","shell.execute_reply.started":"2023-03-08T02:45:37.197208Z","shell.execute_reply":"2023-03-08T02:45:37.208835Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Convert to DataFrame:\n\ndf = pd.DataFrame(fit.history)\ndf","metadata":{"execution":{"iopub.status.busy":"2023-03-08T02:45:37.211729Z","iopub.execute_input":"2023-03-08T02:45:37.212529Z","iopub.status.idle":"2023-03-08T02:45:37.283320Z","shell.execute_reply.started":"2023-03-08T02:45:37.212484Z","shell.execute_reply":"2023-03-08T02:45:37.282158Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Visualization:\n\nplt.figure(figsize = (8, 4))\n\n\n# Graph accuracy:\nplt.subplot(1, 2, 1)\nplt.plot(df['accuracy'])\nplt.plot(df['val_accuracy'])\nplt.xlabel('Epoch')\nplt.ylabel('Accuracy')\nplt.legend(['train', 'validate'], loc = 'lower right');\n\n# Graph loss:\nplt.subplot(1, 2, 2)\nplt.plot(df['loss'])\nplt.plot(df['val_loss'])\nplt.xlabel('Epoch')\nplt.ylabel('Loss')\nplt.legend(['train', 'validate'], loc = 'upper right');","metadata":{"execution":{"iopub.status.busy":"2023-03-08T02:45:37.285304Z","iopub.execute_input":"2023-03-08T02:45:37.285740Z","iopub.status.idle":"2023-03-08T02:45:37.678010Z","shell.execute_reply.started":"2023-03-08T02:45:37.285691Z","shell.execute_reply":"2023-03-08T02:45:37.676703Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Set test dataset path:\n\ntest_data = os.listdir(\"../input/happy-whale-and-dolphin/test_images\")\ntest_data = pd.DataFrame(data = test_data, columns = ['image'])\n\ntest_data['predictions'] = ''\n\ntest_data.head()","metadata":{"execution":{"iopub.status.busy":"2023-03-08T02:45:37.680112Z","iopub.execute_input":"2023-03-08T02:45:37.680956Z","iopub.status.idle":"2023-03-08T02:45:38.205113Z","shell.execute_reply.started":"2023-03-08T02:45:37.680907Z","shell.execute_reply":"2023-03-08T02:45:38.203682Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Loading an image files by its path to predict:\n\ndef Loading_Images(data, length, dataset):\n    X_train = np.zeros((length, 64, 64, 3))\n    count = 0\n    for fig in tqdm(data['image']):\n        img = load_img(\"../input/happy-whale-and-dolphin/\" + dataset + \"/\" + fig, target_size = (64,64, 3))\n        x = img_to_array(img)\n        X_train[count] = x\n        count += 1\n    return X_train\n\n\n# Converting to integers and reshaping:\n\ndef prepare_labels(arr):  \n    values = np.array(arr)\n    label_encoder = LabelEncoder() \n    integer_encoded = label_encoder.fit_transform(values)  # standardization \n    onehot_encoder = OneHotEncoder(sparse = False)         \n    integer_encoded = integer_encoded.reshape(len(integer_encoded), 1) # reshaping\n    onehot_encoded = onehot_encoder.fit_transform(integer_encoded)     # set onehot encoding\n    arr = onehot_encoded\n    return arr, label_encoder","metadata":{"execution":{"iopub.status.busy":"2023-03-08T02:45:38.208062Z","iopub.execute_input":"2023-03-08T02:45:38.208759Z","iopub.status.idle":"2023-03-08T02:45:38.217451Z","shell.execute_reply.started":"2023-03-08T02:45:38.208723Z","shell.execute_reply":"2023-03-08T02:45:38.216638Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Predictions for test dataset:\n\nfrom tqdm.autonotebook import tqdm\nimport gc\n\n#set values:\nele, labels = prepare_labels(train_data['individual_id'])\nlength = len(test_data)\nbatch_size = 1000\nbatch_start = 0\nbatch_end = batch_size\n\n\nwhile batch_start < length:\n    lim = min(batch_end, length)\n    test_batch = test_data.iloc[batch_start:lim]\n    \n    x = Loading_Images(test_batch, test_batch.shape[0], \"test_images\")    # loading images\n    x = x/225\n    \n    pred = mod.predict(np.array(x), verbose = 1)                          # prediction\n    \n    for ele, predict in enumerate(pred):\n        predictions = predict.argsort()[:-5][::-1]\n        idx = -1\n        s = ''; s1 = ''; s2 = ''\n        \n        for i in predictions:\n            idx = idx + 1\n            if predict[i] > 0.5:\n                s1 = s1 + ' ' + labels.inverse_transform(predictions)[idx] \n            else:\n                s2 = s2 + ' ' + labels.inverse_transform(predictions)[idx]\n                \n        s = s1 + 'new_individual' + s2                                    # adding prediction values to test dataset:\n        s = s.strip(' ')\n        \n        test_data.loc[batch_start + ele, 'predictions'] = s\n        \n    batch_start += batch_size   \n    batch_end += batch_size\n    \n    del x\n    del test_batch\n    del pred\n    gc.collect()","metadata":{"execution":{"iopub.status.busy":"2023-03-08T02:45:38.219111Z","iopub.execute_input":"2023-03-08T02:45:38.219446Z","iopub.status.idle":"2023-03-08T03:39:12.105796Z","shell.execute_reply.started":"2023-03-08T02:45:38.219412Z","shell.execute_reply":"2023-03-08T03:39:12.103959Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Results:","metadata":{}},{"cell_type":"code","source":"# Submit results:\n\ntest_data.to_csv('submission.csv', index = False)\ntest_data","metadata":{"execution":{"iopub.status.busy":"2023-03-08T03:39:12.115128Z","iopub.execute_input":"2023-03-08T03:39:12.115610Z","iopub.status.idle":"2023-03-08T03:39:12.337960Z","shell.execute_reply.started":"2023-03-08T03:39:12.115569Z","shell.execute_reply":"2023-03-08T03:39:12.336495Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Conclusion:\n","metadata":{}},{"cell_type":"markdown","source":"The first thing we can see very quickly is the big difference in loss and accuracy in the simple learning model; we see that the model is learning but with only 3 epochs, it gets ~0.7 accuracy and down to ~0.9 loss. \n\nIn our simple model, we used a Conv2D layer as our base layer. We tried different hyperparameters including filters = 64 to 128, strides = (1,1), and kernel_size = (5,5) for better accuracy.\n\nOverall, we are quite happy with our learning model. This study was able to complete its goal by successfully identifying our unique species of dolphins and whales based on images of their fins.\n\n","metadata":{}},{"cell_type":"markdown","source":"#### Thank You","metadata":{}},{"cell_type":"markdown","source":"https://github.com/shriyutha/Whales-and-Dolphins-Identification-CNN-model","metadata":{}}]}