{"cells":[{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true},"cell_type":"code","source":"import numpy as np\nimport pandas as pd\nimport plotly.express as px","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"First let's look into the training metadata.","execution_count":null},{"metadata":{"_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","trusted":true},"cell_type":"code","source":"train = pd.read_csv('/kaggle/input/birdsong-recognition/train.csv')\ntrain.head()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"The quality of a recording is depicted in a rating from 0 to 5.","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"ratings = train['rating'].value_counts().to_frame()\nratings.rename(columns={'rating':'Instances'}, inplace=True)\nratings['Rating'] = ratings.index\nfig = px.bar(ratings, x='Rating', y='Instances',\n            labels={'Instances':'Instances in train'} )\n\nfig.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"The bulk of the data has a rating >3. It might be an idea to leave out recordings with rating <3 when training a model.","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"species = train['species'].value_counts().to_frame()\nfig = px.bar(species, x=species.index, y='species',\n            labels={'species:Species of birds'})\nfig.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"The amount of recordings per bird species show a bias is present. Roughly a third of the data set has got significantly less recordings available in the training data set.","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"channels = train['channels'].value_counts().to_frame()\nchannels.rename(columns={'channels':'Instances'}, inplace=True)\nchannels['Channel'] = channels.index\nchannels.reset_index(drop=True, inplace=True)\nchannels.head()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"The recordings made with a mono or stereo channel are roughly 50/50. Maybe a different model for each of the channels is needed to make accurate predictions.","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"train['month'] = train.date.str[5:7].astype(int)\nmonths = train.month.value_counts().to_frame()\nmonths.rename(columns={'month':'Instances'}, inplace=True)\nmonths['Month'] = months.index\nmonths.reset_index(drop=True, inplace=True)\nmonths.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"fig = px.bar(months, x='Month', y='Instances')\n\nfig.update_layout(\n    title=\"Distribution of bird calls recorded over the months\"\n)\nfig.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"look_up = {1: 'Jan', 2: 'Feb', 3: 'Mar', 4: 'Apr', 5: 'May',\n            6: 'Jun', 7: 'Jul', 8: 'Aug', 9: 'Sep', 10: 'Oct', 11: 'Nov', 12: 'Dec', 0 : '13th'}\nmonths.Month = months.Month.apply(lambda x: look_up[x])\nfig = px.bar(months, x='Month', y='Instances')\nfig.update_layout(\n    title=\"Distribution of bird calls recorded over the months ordered by instances\"\n)\nfig.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Apparently recording birds is more fun in summer. Also 34 birds have been recorded in the 13th month. In all seriousness the data might be biased to species of birds which are active in summer.","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"pitches = train['pitch'].value_counts().to_frame()\npitches.rename(columns={'pitch':'Instances'}, inplace=True)\npitches['pitch'] = pitches.index\npitches.reset_index(drop=True, inplace=True)\npitches.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"fig = px.bar(pitches, x='pitch', y='Instances')\nfig.update_layout(\n    title=\"Distribution of bird call pitch\",\n    xaxis_title=\"Pitch\"\n)\nfig.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"fig = px.box(train, y='duration')\nfig.update_layout(\n    title=\"Duration of recordings\",\n    yaxis_title=\"Duration (seconds)\"\n)\nfig.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Why are some of these recording so long? ","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"times = train.time.value_counts().to_frame()\ntimes.rename(columns={'time':'Instances'}, inplace=True)\ntimes['time'] = times.index\ntimes.reset_index(drop=True, inplace=True)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"fig = px.bar(times, x='time', y='Instances')\nfig.update_layout(\n    title=\"Distribution of bird calls per time stamp\",\n    xaxis_title=\"Pitch\"\n)\nfig.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Time data needs some cleaning up... and it seems a lot of bird recordings are made either at 08:00 or 20:00.","execution_count":null}],"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat":4,"nbformat_minor":4}