{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Competition Introduction\n## Objective\nIn the training set, Individuals have been manually identified and given an individual_id by marine researches. Our task is to individually identify each of the Whales and Dolphins in the test set images using their unique individual characteristics such as shapes, features and markings (some natural, some acquired) of dorsal fins, backs, heads and flanks. This is similar to uniquely identifying human individuals by looking at their faces and other visual features of the body\n\n## Evaluation\nThe evaluation metric is the Mean Average Precision @ 5 (MAP@5):\n$$\nMAP@5 = \\frac{1}{U} \\sum_{u=1}^{U}  \\sum_{k=1}^{min(n,5)} P(k) \\times rel(k)\n$$\nwhere U is the number of customers, P(k) is the precision at cutoff k, n is the number predictions per image, and rel(k) is an indicator function equaling 1 if the item at rank k is a relevant (correct) label, zero otherwise.","metadata":{}},{"cell_type":"markdown","source":"# Exploring the data","metadata":{}},{"cell_type":"markdown","source":"## Exploring train.csv","metadata":{}},{"cell_type":"code","source":"import numpy as np \nimport pandas as pd\nimport os\nimport matplotlib.pyplot as plt\nimport plotly.express as px\nimport cv2","metadata":{"execution":{"iopub.status.busy":"2022-02-12T15:39:13.878821Z","iopub.execute_input":"2022-02-12T15:39:13.879410Z","iopub.status.idle":"2022-02-12T15:39:16.034180Z","shell.execute_reply.started":"2022-02-12T15:39:13.879296Z","shell.execute_reply":"2022-02-12T15:39:16.033140Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df = pd.read_csv('../input/happy-whale-and-dolphin/train.csv')\ntrain_df.head()","metadata":{"execution":{"iopub.status.busy":"2022-02-12T15:39:16.036057Z","iopub.execute_input":"2022-02-12T15:39:16.036304Z","iopub.status.idle":"2022-02-12T15:39:16.154037Z","shell.execute_reply.started":"2022-02-12T15:39:16.036271Z","shell.execute_reply":"2022-02-12T15:39:16.153195Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df.info()","metadata":{"execution":{"iopub.status.busy":"2022-02-12T11:13:28.094141Z","iopub.execute_input":"2022-02-12T11:13:28.094449Z","iopub.status.idle":"2022-02-12T11:13:28.128335Z","shell.execute_reply.started":"2022-02-12T11:13:28.094421Z","shell.execute_reply":"2022-02-12T11:13:28.127613Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(f'There are {train_df[\"individual_id\"].count()} images corresponding to {train_df[\"individual_id\"].nunique()} individuals')","metadata":{"execution":{"iopub.status.busy":"2022-02-12T11:15:26.666578Z","iopub.execute_input":"2022-02-12T11:15:26.667195Z","iopub.status.idle":"2022-02-12T11:15:26.686953Z","shell.execute_reply.started":"2022-02-12T11:15:26.667139Z","shell.execute_reply":"2022-02-12T11:15:26.686199Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sorted(train_df['species'].unique())","metadata":{"execution":{"iopub.status.busy":"2022-02-12T11:08:24.093871Z","iopub.execute_input":"2022-02-12T11:08:24.094421Z","iopub.status.idle":"2022-02-12T11:08:24.101701Z","shell.execute_reply.started":"2022-02-12T11:08:24.094396Z","shell.execute_reply":"2022-02-12T11:08:24.101342Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"As we can see, there are some inconsistent names in the spcies column. Let's correct these using the code snippet shared in [this](https://www.kaggle.com/c/happy-whale-and-dolphin/discussion/305574) discussion thread","metadata":{}},{"cell_type":"code","source":"train_df.species.replace({\"globis\": \"short_finned_pilot_whale\",\n                          \"pilot_whale\": \"short_finned_pilot_whale\",\n                          \"kiler_whale\": \"killer_whale\",\n                          \"bottlenose_dolpin\": \"bottlenose_dolphin\"}, inplace=True)","metadata":{"execution":{"iopub.status.busy":"2022-02-12T11:10:46.828668Z","iopub.execute_input":"2022-02-12T11:10:46.828934Z","iopub.status.idle":"2022-02-12T11:10:46.842388Z","shell.execute_reply.started":"2022-02-12T11:10:46.828912Z","shell.execute_reply":"2022-02-12T11:10:46.841662Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### distribution of train images by species","metadata":{}},{"cell_type":"code","source":"# Now let us plot the distribution of train images by species\nfig = px.histogram(train_df, x=\"species\")\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2022-02-12T11:12:56.868865Z","iopub.execute_input":"2022-02-12T11:12:56.869103Z","iopub.status.idle":"2022-02-12T11:12:58.109345Z","shell.execute_reply.started":"2022-02-12T11:12:56.869079Z","shell.execute_reply":"2022-02-12T11:12:58.108824Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Distribution of train images at individual level\nNow let us see how the distribution of images for each individual level i.e. how many average and maximum images for any individual and see this separately for each species","metadata":{}},{"cell_type":"code","source":"# Images per individual\ndf2 = train_df.groupby(['individual_id','species'])['image'].agg('count').reset_index()\nfig2 = px.box(df2, y=\"image\", color='species', log_y=True, labels={'image':'Number of images'})\nfig2.show()","metadata":{"execution":{"iopub.status.busy":"2022-02-12T16:04:18.484444Z","iopub.execute_input":"2022-02-12T16:04:18.484953Z","iopub.status.idle":"2022-02-12T16:04:18.731269Z","shell.execute_reply.started":"2022-02-12T16:04:18.484914Z","shell.execute_reply":"2022-02-12T16:04:18.730254Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Images\nLet us see a few sample images from the dataset","metadata":{}},{"cell_type":"code","source":"rows, cols = 10,2\nsample = train_df.sample(rows*cols)\nprint(sample)\nf, ax = plt.subplots(rows, cols, figsize=(20,50))\nroot_path = '../input/happy-whale-and-dolphin/train_images/'\nfor i,indx in enumerate(sample.index):\n    file = sample.loc[indx, 'image']\n    species = sample.loc[indx, 'species']\n    individual_id = sample.loc[indx, 'individual_id']\n    img = cv2.imread(root_path+file)\n    title = f'\\n individual_id: {individual_id} \\n Species: {species}'\n    row, col = i//cols, i%cols\n    ax[row, col].imshow(cv2.cvtColor(img, cv2.COLOR_BGR2RGB))\n    ax[row, col].set_title(title)\n    ax[row, col].axis('off')\nplt.tight_layout()\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-02-12T12:48:14.292324Z","iopub.execute_input":"2022-02-12T12:48:14.292622Z","iopub.status.idle":"2022-02-12T12:48:26.479827Z","shell.execute_reply.started":"2022-02-12T12:48:14.292594Z","shell.execute_reply":"2022-02-12T12:48:26.478876Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Following points are can be noted from these sample images:-\n1. Image sizes are not uniform, there are many different image sizes\n2. In addition to image sizes, the distance of the subject individual, and the part of body captured also vary significantly. These are in addition to usual image variables like contrast, focus etc.","metadata":{}},{"cell_type":"markdown","source":"## Datasets of resized images\nKagglers have created and generously shared datasets of resized images which might be good for training initial prototypes of the models. Below are the links to some of these:-\n1. 512 x 512 by phalanx: https://www.kaggle.com/phalanx/whale2-cropped-dataset\n2. 256 x 256 by RDizzl3: https://www.kaggle.com/rdizzl3/jpeg-happywhale-256x256","metadata":{}},{"cell_type":"markdown","source":"## *Work in Progress*\nThis is a work in progress, you are welcome to share your feedback in the comments.","metadata":{}}]}