{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"We can have a first look at the [data](https://www.kaggle.com/c/happy-whale-and-dolphin/data) using the data explorer of Kaggle. We can see that images have different qualities ranging from a dorsal fin to a distant view of the back of the mammal.\n\nLet’s have a more complete view. First add the data through the Kaggle UI. They are afterwards located in */kaggle/input/happy-whale-and-dolphin* folder.","metadata":{"_kg_hide-input":false}},{"cell_type":"code","source":"!ls -l /kaggle/input/happy-whale-and-dolphin","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-03-08T07:45:11.58845Z","iopub.execute_input":"2022-03-08T07:45:11.589795Z","iopub.status.idle":"2022-03-08T07:45:12.441343Z","shell.execute_reply.started":"2022-03-08T07:45:11.589642Z","shell.execute_reply":"2022-03-08T07:45:12.440259Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Data loading\n\nLet’s load the images metadata described in `train.csv` file. \n\nWe can notice that our train data describes 51 033 images with one ID field for the photography filename and 2 others fields. The latter describe the animal specy and which individual it is.\n","metadata":{}},{"cell_type":"code","source":"import pandas as pd\ndf = pd.read_csv(\"/kaggle/input/happy-whale-and-dolphin/train.csv\")\ndf","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2022-03-08T13:55:18.372143Z","iopub.execute_input":"2022-03-08T13:55:18.372411Z","iopub.status.idle":"2022-03-08T13:55:18.444904Z","shell.execute_reply.started":"2022-03-08T13:55:18.372381Z","shell.execute_reply":"2022-03-08T13:55:18.444039Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Data quality\n\n## Field `species`\n\nIn the taxonomy [dataset](https://www.kaggle.com/chasset/happywhalespeciesclassification), authors propose to correct the field `species`.","metadata":{}},{"cell_type":"code","source":"df.loc[df.species == 'bottlenose_dolpin', 'species'] = 'bottlenose_dolphin'\ndf.loc[df.species == 'kiler_whale', 'species'] = 'killer_whale'\ndf.loc[df.species == 'long_finned_pilot_whale', 'species'] = 'pilot_whale'\ndf.loc[df.species == 'globis', 'species'] = 'pilot_whale'\ndf\n","metadata":{"execution":{"iopub.status.busy":"2022-03-08T14:02:46.181520Z","iopub.execute_input":"2022-03-08T14:02:46.181812Z","iopub.status.idle":"2022-03-08T14:02:46.208960Z","shell.execute_reply.started":"2022-03-08T14:02:46.181780Z","shell.execute_reply":"2022-03-08T14:02:46.208204Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df.groupby(['species']).count().sort_values(['species']).shape","metadata":{"execution":{"iopub.status.busy":"2022-03-08T14:02:40.087054Z","iopub.execute_input":"2022-03-08T14:02:40.087722Z","iopub.status.idle":"2022-03-08T14:02:40.106854Z","shell.execute_reply.started":"2022-03-08T14:02:40.087661Z","shell.execute_reply":"2022-03-08T14:02:40.105784Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"taxonomy = pd.read_csv(\"/kaggle/input/happywhalespeciesclassification/species.csv\").set_index('specy_id').drop(columns = ['wikipedia','image','size'])\ntaxonomy","metadata":{"execution":{"iopub.status.busy":"2022-03-08T14:02:15.569744Z","iopub.execute_input":"2022-03-08T14:02:15.570031Z","iopub.status.idle":"2022-03-08T14:02:15.589003Z","shell.execute_reply.started":"2022-03-08T14:02:15.569997Z","shell.execute_reply":"2022-03-08T14:02:15.588158Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Data bias\n\n### Around individuals\n\nThe 51 033 pictures describe 15 887 individuals. The distribution is very skewed:\n\n- 75% of individuals have only one or two pictures\n- 400 pictures are dedicated to only one individual","metadata":{}},{"cell_type":"code","source":"individuals = df.drop(columns = ['species']).groupby(['individual_id']).count().rename(columns = { 'image': 'images_count'})\nindividuals","metadata":{"execution":{"iopub.status.busy":"2022-03-08T07:45:12.576279Z","iopub.execute_input":"2022-03-08T07:45:12.576808Z","iopub.status.idle":"2022-03-08T07:45:12.656111Z","shell.execute_reply.started":"2022-03-08T07:45:12.576775Z","shell.execute_reply":"2022-03-08T07:45:12.654917Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"individuals.describe()","metadata":{"execution":{"iopub.status.busy":"2022-03-08T07:45:12.658598Z","iopub.execute_input":"2022-03-08T07:45:12.658982Z","iopub.status.idle":"2022-03-08T07:45:12.686375Z","shell.execute_reply.started":"2022-03-08T07:45:12.658947Z","shell.execute_reply":"2022-03-08T07:45:12.68547Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from matplotlib import pyplot as plt\nplt.hist(individuals.images_count, density=True, facecolor='g', alpha=0.75)\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-03-08T07:45:12.688141Z","iopub.execute_input":"2022-03-08T07:45:12.688983Z","iopub.status.idle":"2022-03-08T07:45:12.964966Z","shell.execute_reply.started":"2022-03-08T07:45:12.688903Z","shell.execute_reply":"2022-03-08T07:45:12.963775Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Around species\n\nThe 51 887 images describe 30 species. Again, the distribution is skewed. Frasiers dolphin specy is descibed by only 14 images, while Bottle nose dolphin (9664), Beluga (7443) and Humpback whale (7392) are over represented.","metadata":{}},{"cell_type":"code","source":"species = df.drop(columns = ['individual_id']).groupby(['species']).count().rename(columns = { 'image': 'images_count'}).sort_values(by = ['images_count'], ascending = False)\nprint(species.shape)\nspecies","metadata":{"execution":{"iopub.status.busy":"2022-03-08T07:45:12.966545Z","iopub.execute_input":"2022-03-08T07:45:12.966863Z","iopub.status.idle":"2022-03-08T07:45:12.999766Z","shell.execute_reply.started":"2022-03-08T07:45:12.96683Z","shell.execute_reply":"2022-03-08T07:45:12.998835Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"species.describe()","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-03-08T07:45:13.001384Z","iopub.execute_input":"2022-03-08T07:45:13.002273Z","iopub.status.idle":"2022-03-08T07:45:13.017257Z","shell.execute_reply.started":"2022-03-08T07:45:13.002214Z","shell.execute_reply":"2022-03-08T07:45:13.016034Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.hist(species.images_count, density=True, facecolor='g', alpha=0.75)\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-03-08T07:45:13.019098Z","iopub.execute_input":"2022-03-08T07:45:13.01946Z","iopub.status.idle":"2022-03-08T07:45:13.371146Z","shell.execute_reply.started":"2022-03-08T07:45:13.019425Z","shell.execute_reply":"2022-03-08T07:45:13.370137Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Model quality discussion\n\nSo, all together, data contains 51 033 images of 15 587 individuals from 30 species. But few species have a good ratio. This will affect the prediction quality. Many species/individuals will be difficult to predict.","metadata":{}},{"cell_type":"code","source":"individuals_counts = df.drop(columns = ['image']).groupby(['species']).nunique().rename(columns = { 'individual_id': 'individuals_count'})\ncounts = pd.merge(species, individuals_counts, how = 'left', on = ['species'])\ncounts = counts.assign(ratio = counts.images_count / counts.individuals_count).sort_values(by = ['ratio'], ascending = False)\ncounts","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-03-08T07:45:13.372388Z","iopub.execute_input":"2022-03-08T07:45:13.372625Z","iopub.status.idle":"2022-03-08T07:45:13.430084Z","shell.execute_reply.started":"2022-03-08T07:45:13.372598Z","shell.execute_reply":"2022-03-08T07:45:13.42892Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.hist(counts.ratio, density=True, facecolor='g', alpha=0.75)\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-03-08T07:45:13.4321Z","iopub.execute_input":"2022-03-08T07:45:13.43242Z","iopub.status.idle":"2022-03-08T07:45:13.660979Z","shell.execute_reply.started":"2022-03-08T07:45:13.432386Z","shell.execute_reply":"2022-03-08T07:45:13.659804Z"},"trusted":true},"execution_count":null,"outputs":[]}]}