{"cells":[{"metadata":{},"cell_type":"markdown","source":"# EDA of Meta Data","execution_count":null},{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true},"cell_type":"code","source":"import os\nimport numpy as np\nimport pandas as pd\nimport librosa\nimport librosa.display\nimport plotly.express as px\nimport plotly.graph_objects as go\nimport folium\nimport IPython\nimport matplotlib.pyplot as plt","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","trusted":true},"cell_type":"code","source":"print('\\n'.join(os.listdir('/kaggle/input/birdsong-recognition')))\npath = '/kaggle/input/birdsong-recognition'\nTRAIN_PATH = os.path.join(path, 'train_audio')\nTEST_PATH = os.path.join(path, 'example_test_audio')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train_meta = pd.read_csv(os.path.join(path, 'train.csv'))\ntest_meta = pd.read_csv(os.path.join(path, 'test.csv'))","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Rating\nThe rating of the audio of the quality of the audio file and the confidence of the label on the birdcall. This depends on if the researcher is able to see the bird or the amount of ambient noise present. A lower rating means a softer label.","execution_count":null},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"rating_count = train_meta.groupby('rating')['xc_id'].count().reset_index()\nrating_count.rename(columns={'xc_id': 'Count', 'rating': 'Rating'}, inplace=True)\nfig = px.bar(rating_count, x='Rating', y='Count', title='Bar Plot of Number of Audio Recordings in Each Rating Category')\nfig.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Distribution of Labels and Training Data","execution_count":null},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"labels = train_meta.groupby(['species'])['xc_id'].count().reset_index()\nlabels.rename(columns={'xc_id': 'Count', 'species': 'Species'}, inplace=True)\nlabels.sort_values(by=['Count'], inplace=True, ascending=True)\nfig = px.bar(labels, x='Species', y='Count', hover_data=['Species'], color='Species')\nfig.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"We can see that most classes have almost the exact same amount of data: 100 files. However, for a minority of classes there is data imbalance present. This could possible result in underfitting in these categories; namely, `Redhead`, `Bufflehead`, etc.","execution_count":null},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"fig = px.histogram(train_meta['duration'].reset_index(), x='duration', labels={'count': 'Count', 'duration': 'Duration'},\n                   title='Distribution of Duration of Audio Files')\nfig.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"As the duration of audio file increases there is more probability that there will be ambient noise. If we train a model by splitting the audio files into 5 second partitions there is a possibility that some 5 second partitions will just be ambient noise. ","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"## Sampling Rate","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"train_meta.sampling_rate = train_meta.sampling_rate.apply(lambda sr: int(sr.split(' ')[0])).astype(np.uint16)\nfig = px.histogram(train_meta['sampling_rate'].reset_index(), x='sampling_rate', labels={'count': 'Count', 'sampling_rate': 'Sampling Rate'},\n                   title='Distribution of Sampling Rate of Audio Files')\nfig.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Sample Breakdown","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"sample1 = train_meta.iloc[528]","execution_count":null,"outputs":[]},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"IPython.display.Audio(filename=os.path.join(TRAIN_PATH, \n                                            os.path.join(sample1['ebird_code'],sample1['filename'])))","execution_count":null,"outputs":[]},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"sample1","execution_count":null,"outputs":[]},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"y, sr = librosa.load(os.path.join(os.path.join(TRAIN_PATH, sample1['ebird_code']), sample1['filename']), \n                     sr=sample1['sampling_rate'])\nsample1_audio, _ = librosa.effects.trim(y)\ntime = [v/sample1['sampling_rate'] for v in range(len(sample1_audio))]\nfig = px.line({'Time': time, 'Intensity': sample1_audio}, x='Time', y='Intensity', title='Wave Plot of Sample1 Audio')\nfig.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Let's see what the frequency domain looks like after we apply a fourier transformation. A fourier transformation takes an audio clip that is in the time domain and transforms it into the frequency domain. Here is a great video on [it](https://www.youtube.com/watch?v=spUNpyF58BY). <---","execution_count":null},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"n_fft = 2048\nD = np.abs(librosa.stft(sample1_audio[:n_fft], n_fft=n_fft, hop_length=n_fft+1))\nfig = px.line({'Frequency': np.array(range(len(D))), 'Magnitude': D.flatten()}, x='Frequency', y='Magnitude', \n              title='Fourier Transformation of Sample1 Audio')\nfig.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Spectrogram of Sample 1","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"hop_length = 512\nn_fft = 2048\nD = np.abs(librosa.stft(sample1_audio, n_fft=n_fft,  hop_length=hop_length))\nlibrosa.display.specshow(D, sr=sr, x_axis='time', y_axis='linear');\nplt.colorbar();","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"It's really hard to see anything because the sounds that most humans hear are concentrated in the lower frequencies. We can adjust the intensity to decibles which is the log-scale for intensity; we can also use the [mel-scale](https://towardsdatascience.com/getting-to-know-the-mel-spectrogram-31bca3e2d9d0) to represent frequencies. ","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"DB = librosa.amplitude_to_db(D, ref=np.max)\nlibrosa.display.specshow(DB, sr=sr, hop_length=hop_length, x_axis='time', y_axis='log');\nplt.colorbar(format='%+2.0f dB');","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"This is our spectrogram ! By using spectrogram representations of our audio, we can perform image classification on these spectrograms. This makes it easy to use pre-existing DNN architectures that have been tried and proven.","execution_count":null}],"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat":4,"nbformat_minor":4}