{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Incorrect spelling in the Label\n\n## Among multiple categories of labels present in the data, a specific category seems to be misspelled. Below are the 2 categories:\n* bottlenose_dolphin\n* bottlenose_dolpin --> misspelled\n\nI think this is a case of misspelling in the label as nothing called \"dolpin\" exists in english dictionary. This case of misspelling causes a data split anong the same type of dolphins, during training phase the model would get confused when same type of images of dolphin species will be presented but will be forced to learn them as different species although that is not the case. This will eventually hurt the testing results as well as the model will predict images as\"bottlenose-dolpin\" which may not exist in the test image cases.\n\nTo observe the evidence of this misspelled class and its solution please see below:","metadata":{}},{"cell_type":"code","source":"import numpy as np\nimport pandas as pd \nimport os","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"!ls /kaggle/input/happy-whale-and-dolphin","metadata":{"execution":{"iopub.status.busy":"2022-04-10T19:12:34.395403Z","iopub.execute_input":"2022-04-10T19:12:34.396298Z","iopub.status.idle":"2022-04-10T19:12:35.154323Z","shell.execute_reply.started":"2022-04-10T19:12:34.396249Z","shell.execute_reply":"2022-04-10T19:12:35.153127Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"inputPath = '/kaggle/input/happy-whale-and-dolphin'","metadata":{"execution":{"iopub.status.busy":"2022-04-10T19:12:35.158169Z","iopub.execute_input":"2022-04-10T19:12:35.158923Z","iopub.status.idle":"2022-04-10T19:12:35.163998Z","shell.execute_reply.started":"2022-04-10T19:12:35.158876Z","shell.execute_reply":"2022-04-10T19:12:35.163327Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df = pd.read_csv(inputPath+'/train.csv')","metadata":{"execution":{"iopub.status.busy":"2022-04-10T19:12:36.126449Z","iopub.execute_input":"2022-04-10T19:12:36.126775Z","iopub.status.idle":"2022-04-10T19:12:36.242993Z","shell.execute_reply.started":"2022-04-10T19:12:36.126744Z","shell.execute_reply":"2022-04-10T19:12:36.241499Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df.head(5)","metadata":{"execution":{"iopub.status.busy":"2022-04-10T19:12:48.669306Z","iopub.execute_input":"2022-04-10T19:12:48.670225Z","iopub.status.idle":"2022-04-10T19:12:48.696135Z","shell.execute_reply.started":"2022-04-10T19:12:48.670175Z","shell.execute_reply":"2022-04-10T19:12:48.695186Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df['species'].value_counts().plot(kind = 'bar')","metadata":{"execution":{"iopub.status.busy":"2022-04-10T19:13:49.634954Z","iopub.execute_input":"2022-04-10T19:13:49.635626Z","iopub.status.idle":"2022-04-10T19:13:50.353537Z","shell.execute_reply.started":"2022-04-10T19:13:49.635567Z","shell.execute_reply":"2022-04-10T19:13:50.352530Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df['family'] = train_df['species'].str.split('_').str[-1]","metadata":{"execution":{"iopub.status.busy":"2022-04-10T19:17:37.989123Z","iopub.execute_input":"2022-04-10T19:17:37.989466Z","iopub.status.idle":"2022-04-10T19:17:38.090375Z","shell.execute_reply.started":"2022-04-10T19:17:37.989430Z","shell.execute_reply":"2022-04-10T19:17:38.089204Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df['family'].value_counts().plot(kind = 'bar')","metadata":{"execution":{"iopub.status.busy":"2022-04-10T19:17:49.203934Z","iopub.execute_input":"2022-04-10T19:17:49.204313Z","iopub.status.idle":"2022-04-10T19:17:49.429749Z","shell.execute_reply.started":"2022-04-10T19:17:49.204271Z","shell.execute_reply":"2022-04-10T19:17:49.428558Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df['family'].value_counts()","metadata":{"execution":{"iopub.status.busy":"2022-04-10T19:20:28.863625Z","iopub.execute_input":"2022-04-10T19:20:28.863990Z","iopub.status.idle":"2022-04-10T19:20:28.883814Z","shell.execute_reply.started":"2022-04-10T19:20:28.863954Z","shell.execute_reply":"2022-04-10T19:20:28.882784Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df[train_df['species'].str.contains('beluga')]['species'].unique()","metadata":{"execution":{"iopub.status.busy":"2022-04-10T19:24:36.456295Z","iopub.execute_input":"2022-04-10T19:24:36.456659Z","iopub.status.idle":"2022-04-10T19:24:36.496835Z","shell.execute_reply.started":"2022-04-10T19:24:36.456620Z","shell.execute_reply":"2022-04-10T19:24:36.495727Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df[train_df['species'].str.contains('whale')]['species'].unique()","metadata":{"execution":{"iopub.status.busy":"2022-04-10T19:25:20.783065Z","iopub.execute_input":"2022-04-10T19:25:20.783456Z","iopub.status.idle":"2022-04-10T19:25:20.830967Z","shell.execute_reply.started":"2022-04-10T19:25:20.783420Z","shell.execute_reply":"2022-04-10T19:25:20.829947Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Seems like there are no \"Beluga whales\" but a category \"beluga\" exists. That is fine since the models are not presented with any ambiguity as \"Beluga\" and \"beluga_whales\" doesn't make any difference for the models","metadata":{}},{"cell_type":"code","source":"train_df[train_df['species'].str.contains('dolpin')]['species'].unique()","metadata":{"execution":{"iopub.status.busy":"2022-04-10T19:24:49.289579Z","iopub.execute_input":"2022-04-10T19:24:49.290156Z","iopub.status.idle":"2022-04-10T19:24:49.329480Z","shell.execute_reply.started":"2022-04-10T19:24:49.290087Z","shell.execute_reply":"2022-04-10T19:24:49.328377Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df[train_df['species'].str.contains('dolphin')]['species'].unique()","metadata":{"execution":{"iopub.status.busy":"2022-04-10T19:25:03.026210Z","iopub.execute_input":"2022-04-10T19:25:03.026742Z","iopub.status.idle":"2022-04-10T19:25:03.070403Z","shell.execute_reply.started":"2022-04-10T19:25:03.026703Z","shell.execute_reply":"2022-04-10T19:25:03.069339Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## In this case there are 2 categories:\n* bottlenose_dolphin\n* bottlenose_dolpin\n\nThis are the 2 same categories but the difference in name is caused due to incorrect spelling. \n\n## This can be fixed by the following code snippet:","metadata":{}},{"cell_type":"code","source":"train_df['species'] = train_df['species'].str.replace('dolpin','dolphin')","metadata":{"execution":{"iopub.status.busy":"2022-04-10T19:43:03.875051Z","iopub.execute_input":"2022-04-10T19:43:03.875476Z","iopub.status.idle":"2022-04-10T19:43:03.922550Z","shell.execute_reply.started":"2022-04-10T19:43:03.875429Z","shell.execute_reply":"2022-04-10T19:43:03.921661Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df[train_df['species'].str.contains('dolpin')]['species'].unique()","metadata":{"execution":{"iopub.status.busy":"2022-04-10T19:43:15.758624Z","iopub.execute_input":"2022-04-10T19:43:15.759603Z","iopub.status.idle":"2022-04-10T19:43:15.804384Z","shell.execute_reply.started":"2022-04-10T19:43:15.759540Z","shell.execute_reply":"2022-04-10T19:43:15.803659Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"","metadata":{}}]}