{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Packages used","metadata":{}},{"cell_type":"code","source":"import numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\nimport matplotlib.pyplot as plt","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Loading the training dataframe","metadata":{}},{"cell_type":"code","source":"train_df = pd.read_csv('../input/happy-whale-and-dolphin/train.csv')","metadata":{"execution":{"iopub.status.busy":"2022-03-09T01:38:28.268009Z","iopub.execute_input":"2022-03-09T01:38:28.268415Z","iopub.status.idle":"2022-03-09T01:38:28.382274Z","shell.execute_reply.started":"2022-03-09T01:38:28.268374Z","shell.execute_reply":"2022-03-09T01:38:28.381133Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df.head()","metadata":{"execution":{"iopub.status.busy":"2022-03-09T01:38:28.383496Z","iopub.execute_input":"2022-03-09T01:38:28.383735Z","iopub.status.idle":"2022-03-09T01:38:28.404415Z","shell.execute_reply.started":"2022-03-09T01:38:28.383704Z","shell.execute_reply":"2022-03-09T01:38:28.403548Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# General information on species and individuals","metadata":{}},{"cell_type":"markdown","source":"To start, let's retrieve the species names and how many of each there are in the dataset.","metadata":{}},{"cell_type":"code","source":"train_df.species.groupby(train_df.species).count()","metadata":{"execution":{"iopub.status.busy":"2022-03-09T01:38:28.405690Z","iopub.execute_input":"2022-03-09T01:38:28.405928Z","iopub.status.idle":"2022-03-09T01:38:28.431881Z","shell.execute_reply.started":"2022-03-09T01:38:28.405898Z","shell.execute_reply":"2022-03-09T01:38:28.431325Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"num_species = train_df.species.groupby(train_df.species).count().nunique()\nprint('Number of species:', num_species)","metadata":{"execution":{"iopub.status.busy":"2022-03-09T01:38:28.433072Z","iopub.execute_input":"2022-03-09T01:38:28.433303Z","iopub.status.idle":"2022-03-09T01:38:28.451673Z","shell.execute_reply.started":"2022-03-09T01:38:28.433276Z","shell.execute_reply":"2022-03-09T01:38:28.450840Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"As we can see above, there are some misspelled species names. Therefore, a replacement must be taken into account.\nThe issue was also identified by another Kaggler, [Aleksey Alekssev](https://www.kaggle.com/c/happy-whale-and-dolphin/discussion/305574).","metadata":{}},{"cell_type":"code","source":"# from Aleksey Alekssev\nreplacement_data = {'globis': 'short_finned_pilot_whale',\n                    'pilot_whale': 'short_finned_pilot_whale',\n                    'kiler_whale': 'killer_whale',\n                    'bottlenose_dolpin': 'bottlenose_dolphin'}\n\ntrain_df.species.replace(replacement_data, inplace=True)\n\ntrain_df.species.groupby(train_df.species).count()","metadata":{"execution":{"iopub.status.busy":"2022-03-09T01:38:28.453237Z","iopub.execute_input":"2022-03-09T01:38:28.453848Z","iopub.status.idle":"2022-03-09T01:38:28.492351Z","shell.execute_reply.started":"2022-03-09T01:38:28.453799Z","shell.execute_reply":"2022-03-09T01:38:28.490527Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"To have a visualization on how the species are distributed, we can make use of a histogram.","metadata":{}},{"cell_type":"code","source":"plt.figure(figsize=(10, 6))\ntrain_df.species.value_counts().sort_values(ascending=True).plot(kind='barh')","metadata":{"execution":{"iopub.status.busy":"2022-03-09T01:38:28.514508Z","iopub.execute_input":"2022-03-09T01:38:28.514766Z","iopub.status.idle":"2022-03-09T01:38:28.949673Z","shell.execute_reply.started":"2022-03-09T01:38:28.514735Z","shell.execute_reply":"2022-03-09T01:38:28.948836Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"And finally, some basic descriptive statistics about the `species` and `individual_id`.","metadata":{}},{"cell_type":"code","source":"train_df[['species', 'individual_id']].describe()","metadata":{"execution":{"iopub.status.busy":"2022-03-09T02:09:03.124169Z","iopub.execute_input":"2022-03-09T02:09:03.124442Z","iopub.status.idle":"2022-03-09T02:09:03.172677Z","shell.execute_reply.started":"2022-03-09T02:09:03.124413Z","shell.execute_reply":"2022-03-09T02:09:03.171832Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"From all the information retrieved so far, we now know how many classes should be used in our CNN model.\n* 15587 classes for the known unique individuals plus 1 class for never seen individuals.","metadata":{}},{"cell_type":"markdown","source":"As a complementary information, let's see who is the most famous individual in the dataset.","metadata":{}},{"cell_type":"code","source":"train_df[train_df.individual_id == '37c7aba965a5'].head(1)","metadata":{"execution":{"iopub.status.busy":"2022-03-09T02:11:42.582200Z","iopub.execute_input":"2022-03-09T02:11:42.582847Z","iopub.status.idle":"2022-03-09T02:11:42.604586Z","shell.execute_reply.started":"2022-03-09T02:11:42.582812Z","shell.execute_reply":"2022-03-09T02:11:42.603880Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Who else appears many times in the dataset?","metadata":{}},{"cell_type":"code","source":"train_df.individual_id.value_counts().sort_values(ascending=False).head(10)","metadata":{"execution":{"iopub.status.busy":"2022-03-09T02:24:19.251665Z","iopub.execute_input":"2022-03-09T02:24:19.252154Z","iopub.status.idle":"2022-03-09T02:24:19.274969Z","shell.execute_reply.started":"2022-03-09T02:24:19.252121Z","shell.execute_reply":"2022-03-09T02:24:19.274107Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Taking into account that I am thinking about using CNN to tackle this competition problem, and the fact that from `individual_id` we can see that our training data is actually unballaced, in this case, I wonder about the following:\n> To work with a CNN, should such unbalaced data be an issue for the final model?\n\nAnd regarding the images size:\n> What could be a good initial approach to deal if different images sizes in order to feed them to the CNN?\n\nkaggler RDizzl3 shared a topic on [Reduced Resolution Image Data](https://www.kaggle.com/c/happy-whale-and-dolphin/discussion/304686) to improve computational performance for those working for the first time with computational vision.\n\n\n\n\n","metadata":{}}]}