{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Set up\nTo the right, there's a \"Settings\" tab. Go there and turn the internet on. After that, run the next code block, and install the fastaudio library.\n\nI use it because it makes things easy to run!","metadata":{}},{"cell_type":"code","source":"!pip install --upgrade fastaudio","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\nimport os\nimport IPython.display as ipd\nimport librosa\nimport librosa.display\nimport librosa.effects\n\nimport matplotlib.pyplot as plt\nfrom scipy.io import wavfile as wav\n\nfrom sklearn import metrics \nfrom sklearn.preprocessing import LabelEncoder\nfrom sklearn.model_selection import train_test_split\n\nimport fastaudio.all\nimport fastai.vision.all\n\nfrom tqdm.auto import tqdm\n\n# import torchaudio\n# import torchaudio.functional as F\n# import torchaudio.transforms as T\n#warnings.filterwarnings('ignore')\n#device = torch.device(\"cuda:0\" if torch.cuda.is_available() else \"cpu\")","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Inspiration sources:\n(AKA I kinda ripped them off because I don't know audio processing)\n\n* https://opensource.com/article/19/9/audio-processing-machine-learning-python\n* https://www.kaggle.com/code/utcarshagrawal/birdclef-audio-pytorch-tutorial\n* https://towardsdatascience.com/how-to-apply-machine-learning-and-deep-learning-methods-to-audio-analysis-615e286fcbbc\n* https://www.kaggle.com/code/shreyasajal/birdclef-librosa-audio-feature-extraction","metadata":{}},{"cell_type":"markdown","source":"# Intro to problem\n\n(This is just meant to be a simple introduction, the more detailed information can be found at the competition site [here](https://www.kaggle.com/competitions/birdclef-2022/overview/description))\n\nIt is important to monitor birds species that are in danger of extinction, but due to their habitats it is usually rather hard to monitor them directly. Sound recordings have been found as an alternative, but it is a challenge to use this for endangered species. The objective then is to process audio to recognize a given species.\n\n## Why is this type of ML problem more difficult?\n\nThis is a classification problem, the difficult part of this lies in the fact that we are working with audio files. \n\nThe main difficulty for this kind of problem lies in the processing of audio files. For most ML problems, the input size is fixed. Here, the audio file can be longer or shorter.\n\nOn top of that, most ML programs work only with numerical data. Sound is not really numerical, so something needs to be done with it beforehand. Here in this notebook, we will give some basic ideas on how to do this.\n\n## What is the focus going to be\n\nThere are multiple ways of going around for this, but because I'm not personally familiar with how audio processing works, I'll focus on explaining existing libraries and trying to make the most out of them.\n\nThis problem falls into time a time series model. This is due to how audio work, so that does mean part of the general work will involve time series analysis.\n\nThe two main tools we'll use are librosa and pytorch (focusing on the audio side of things).","metadata":{}},{"cell_type":"markdown","source":"# Importing files\nFirst step will be to have a system that can import files. Because the information is stored in audio files, we need to be able to open the files dynamically according to a file descriptor.\n\nDon't worry, the first step is easier and just involves reading in the csv files already contained.\n\nWe'll take a lazy approach and just visualize each of the csv files inside kaggle directly.\n\nIf there are any relevant insights they'll be mentioned as we go along.","metadata":{}},{"cell_type":"code","source":"submission_sample = pd.read_csv('/kaggle/input/birdclef-2022/sample_submission.csv')\nscored_birds = pd.read_csv('/kaggle/input/birdclef-2022/scored_birds.json')\nebird_taxonomy = pd.read_csv('/kaggle/input/birdclef-2022/eBird_Taxonomy_v2021.csv')\ntest_data = pd.read_csv('/kaggle/input/birdclef-2022/test.csv')\ntrain_data = pd.read_csv('/kaggle/input/birdclef-2022/train_metadata.csv')","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"scored_birds","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"ebird_taxonomy","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"len(ebird_taxonomy['SCI_NAME'].unique())","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"ebird_taxonomy['CATEGORY'].unique()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Each entry provides information on birds. It varies, sometimes it's specific species information, sometimes it's a misc, but there is some information on some birds on each row (very helpful, I know).\n\nSo, the SCI_NAME works as a kind of unique key, each row is distinguished by just that.","metadata":{}},{"cell_type":"code","source":"test_data","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"submission_sample","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"With test data and the submission sample, you can visualize that what you're predicting is whether it's true or false that the audio recording belongs to a specific bird. 'row_id' is the name of the sound file, so that means there is only 3 submissions we can predict on (for now).","metadata":{}},{"cell_type":"code","source":"train_data.head()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"There is plenty of row info in the training data. In theory if you want to be smart you might be able to find relations between species, so for example if two birds are of the same ORDER1, then maybe their sound might be related.\n\nFor now, we are going to be reallyyyyy lazy and just focus on the sound processing. This means we only care about the filename part of things.","metadata":{}},{"cell_type":"code","source":"filename_sample = train_data['filename'].values[0]\n# We're using some ipython magic things so we can hear an audio file\n# It's not really needed, but it makes it much easier to help hear what we're working with\nipd.Audio(f\"../input/birdclef-2022/train_audio/{filename_sample}\")","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Plotting a spectogram\n\nSpectograms are apparently super helpful for processing audio files. Don't know why, but we'll just go along with it.","metadata":{}},{"cell_type":"code","source":"def get_audio_features(file_name: str):\n    audio, sample_rate = librosa.load(file_name)\n    mfccs = librosa.feature.mfcc(y=audio, sr=sample_rate)\n    \n    return \n","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Creating usable inputs\nSo, we apparently have to go first to a spectrum and then transform that into something nicer. This is the preprocessing step, after that we have something that should be easier to do some ML learning on.\n\nWe are going to use librosa for this, because it was recommended online.","metadata":{}},{"cell_type":"code","source":"filename_sample = train_data['filename'].values[0]\nfilename_sample = f\"../input/birdclef-2022/train_audio/{filename_sample}\"\n# We're using some ipython magic things so we can hear an audio file\n# It's not really needed, but it makes it much easier to help hear what we're working with\nlibrosa_audio, librosa_sr = librosa.load(filename_sample)\nprint(\"Sample rate: {}\".format(librosa_sr))\nprint('Audio as a time series shape', np.shape(librosa_audio))","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Librosa also converts audio to Mono. You can visualize it by the graph.\n\n(I wanted to compare with scipy, but scipy doesn't support .oggs, sorry!)","metadata":{"execution":{"iopub.status.busy":"2022-04-05T17:45:38.157144Z","iopub.execute_input":"2022-04-05T17:45:38.157521Z","iopub.status.idle":"2022-04-05T17:45:38.164306Z","shell.execute_reply.started":"2022-04-05T17:45:38.157483Z","shell.execute_reply":"2022-04-05T17:45:38.162892Z"}}},{"cell_type":"code","source":"librosa.display.waveplot(y=librosa_audio, sr= librosa_sr)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Spectogram\n\nA representation of the spectrum of frequencies.","metadata":{}},{"cell_type":"code","source":"stft_lib = np.abs(librosa.stft(librosa_audio))\n# Convert an amplitude spectrogram to Decibels-scaled spectrogram.\ndb_audio = librosa.amplitude_to_db(stft_lib)\n\nlibrosa.display.specshow(db_audio)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# RMSE\nA way of calculating how the signal is (due to how it's related to the energy of the signal)\n","metadata":{}},{"cell_type":"code","source":"S, phase = librosa.magphase(stft_lib)\nS_db=librosa.amplitude_to_db(S)\nrms = librosa.feature.rms(S=S)\ntimes = librosa.times_like(rms)\n\nfig, ax = plt.subplots(nrows=1, sharex=True,figsize = (16, 6))\n\nax.semilogy(times, rms[0], label='RMS Energy')\nax.set(xticks=[])\nax.legend()\nax.label_outer()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Mel-frequency cepstrum\n\nTODO: Fill the information here about what they are\n\nProbably some information to be found from [practicalcryptography](http://practicalcryptography.com/miscellaneous/machine-learning/guide-mel-frequency-cepstral-coefficients-mfccs/) on MFCC\n","metadata":{}},{"cell_type":"code","source":"mfccs = librosa.feature.mfcc(y=librosa_audio, sr=librosa_sr)\nprint(mfccs.shape)\nlibrosa.display.specshow(mfccs, sr=librosa_sr, x_axis='time')","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"np.mean(mfccs.T, axis=0)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We're getting the mean to make it easier to get features. (some of this information should make sense from understanding mcc I guess?)","metadata":{}},{"cell_type":"markdown","source":"For some of the files, there's a small leading and trailing silence. We're going to remove that","metadata":{}},{"cell_type":"code","source":"audio_trimmed, _ = librosa.effects.trim(librosa_audio)\nprint('Trimmed shape:', np.shape(audio_trimmed))\nprint('Default shape:', np.shape(librosa_audio))","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Separating harmonics and percussion\n","metadata":{}},{"cell_type":"code","source":"y_harm, y_perc = librosa.effects.hpss(librosa_audio)\nD_vals = np.abs(librosa.stft(librosa_audio))\nDB_audio = librosa.amplitude_to_db(D_vals)\nplt.figure(figsize = (16, 6))\nplt.plot(y_perc, color = '#FFB100')\nplt.plot(y_harm, color = '#A300F9')\nplt.legend((\"Perceptrual\", \"Harmonics\"))\nplt.title(\"Harmonics + Percussive :\", fontsize=16);\n\n\nH, P = librosa.decompose.hpss(librosa.stft(librosa_audio))    \nplt.figure(figsize=(16, 6))\nplt.subplot(3, 1, 1)\nlibrosa.display.specshow(DB_audio)\nplt.colorbar(format='%+2.0f dB')\nplt.title('Full power spectrogram: Harmonic + Percussive')\n\n# harmonic spectrogram will show more horizontal/pitch-dependent changes\nplt.subplot(3, 1, 2)\nlibrosa.display.specshow(librosa.amplitude_to_db(np.abs(H), ref=np.max))\nplt.colorbar(format='%+2.0f dB')\nplt.title('Harmonic power spectrogram')\nplt.subplot(3, 1, 3)\n\n# percussive spectrogram will show more vertical/time-dependent changes\nlibrosa.display.specshow(librosa.amplitude_to_db(np.abs(P), ref=np.max))\nplt.colorbar(format='%+2.0f dB')\nplt.title('Percussive power spectrogram')\nplt.tight_layout()\nplt.show()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"You can do more cool fancy graphs (this notebook is really good for ideas: https://www.kaggle.com/code/shreyasajal/birdclef-librosa-audio-feature-extraction)\n\nBut we want to do something we can model, so let's start with that!","metadata":{}},{"cell_type":"markdown","source":"# Comparing birds\nGreatly based on the work done here: https://www.kaggle.com/code/ohseokkim/birdclef-2022-shut-up-and-calculate","metadata":{}},{"cell_type":"code","source":"def show_bird(audios):\n    for fn in audios:\n        audio = fastaudio.all.AudioTensor.create(fn)\n        audio.show()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"normoc_fns = fastaudio.all.get_audio_files( '../input/birdclef-2022/train_audio/normoc')\nprint(len(normoc_fns))\n\n# We'll just take the first three for the sake of comparison\nshow_bird(normoc_fns[:3])","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"(If I understand correctly and if my ears work) The difference in channels is due to first and third files being stereo, while the second one is mono.\n\nI don't know much about birds, but none of those calls sounded similar! And there doesn't seem to be much overlap between the waveforms.\n\nAlso, there's some noise in the background. That's definitely not helpful!","metadata":{}},{"cell_type":"markdown","source":"Now, processing audio is kind of hard. But image processing pictures are more explored, and there are existing models that do well. How do we do this?\n\nWe use the mel-frequency cepstrum (MFCC) we visualized earlier. Even though what it contains is not that easy to see from a human eye, computers tend to actually learn from that. So we'll create a way of transforming the audio to MFCC, and then process that data.\n\nWe'll create a pipeline for this process.","metadata":{}},{"cell_type":"code","source":"# I input a value here, because if not the system would give a warning\naud2mfcc = fastaudio.all.AudioToMFCC(n_mfcc=40)\n\nfunctions_pipeline = [fastaudio.all.RemoveSilence(), fastaudio.all.ResizeSignal(1000), aud2mfcc, fastaudio.all.Delta()]","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# In the notebook the import makes things easier by using *\n# I decided to not do it like that, because that way we can kind of visualize what is audio and what is visual processing\naud_digit = fastai.vision.all.DataBlock(blocks=(fastaudio.all.AudioBlock, fastai.vision.all.CategoryBlock),\n                                    get_items= fastaudio.all.get_audio_files,\n                                    splitter = fastai.vision.all.RandomSplitter(),\n                                    item_tfms = functions_pipeline,\n                                    get_y = fastai.vision.all.parent_label)\n\ndata_dir = fastaudio.all.Path('../input/birdclef-2022/train_audio')\naud_digit.summary(data_dir)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"dls = aud_digit.dataloaders(data_dir)\ndls.c","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"That's the number of target classes, so 152.","metadata":{}},{"cell_type":"code","source":"dls.show_batch()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"This is an example of the converted audio. You can see there's a lable at the top for each of the three channels.\n\nSo now the problem is one of image classification.","metadata":{}},{"cell_type":"code","source":"def audio_learner(dls, arch, loss_func, metrics):\n    \"Prepares a `Learner` for audio processing\"\n    learn = fastai.vision.all.Learner(dls, arch, loss_func, metrics=metrics, \n                  cbs = [fastai.vision.all.EarlyStoppingCallback(monitor='accuracy', patience=5),fastai.vision.all.ActivationStats(with_hist=True)]).to_fp16()\n    n_c = dls.one_batch()[0].shape[1]\n    if n_c == 1: alter_learner(learn)\n    return learn\n\nlearn = audio_learner(dls, \n                      fastai.vision.all.xresnet18(), \n                      fastai.vision.all.LabelSmoothingCrossEntropy(), \n                      fastai.vision.all.accuracy)\n\n# Note how we're now mainly using the fastai.vision methods, as we're focusing on image processing","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"learn.fit_one_cycle(1)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"interp = fastai.vision.all.ClassificationInterpretation.from_learner(learn)\ninterp.plot_confusion_matrix(figsize=(30,30), dpi=240)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"That was eh. In this case the graph is way too large to be readable, but our \"focus\" is on the main diagonal. You can see that it is kind of empty, which isn't really the best overall.\n\n\nBut, that's because I was deliberately lazy and didn't focus on optimizing the learning rate!\n\nLearning rate is the step size of how much is learnt at each step. There's a tradeoff, a too low learning rate won't reach a good value in reasonable time, while one that's too high will overshoot too much and never reach the actual optima.","metadata":{}},{"cell_type":"code","source":"sr = learn.lr_find()\nsr","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"learn.fit_one_cycle(1, sr.lr_steep)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def plot_layer_stats(self, idx):\n    plot,axs = plt.subplots(1, 3, figsize=(15,3))\n    plot.subplots_adjust(wspace=0.5)\n    for o,ax,title in zip(self.layer_stats(idx),axs,('mean','std','% near zero')):\n        ax.plot(o)\n        ax.set_title(title)\n        \nplot_layer_stats(learn.activation_stats,-2)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plot_layer_stats(learn.activation_stats,-1)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"learn.recorder.plot_loss()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"interp = fastai.vision.all.ClassificationInterpretation.from_learner(learn)\ninterp.plot_confusion_matrix(figsize=(30,30), dpi=240)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"interp.most_confused(min_val=10)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Still kinda iffy, and there's a lot of confusion.\n\nBut given we've limited training time, this is ok (you can do better!)","metadata":{}},{"cell_type":"markdown","source":"# Doing a submission\nThis dataset is a bit of a pain with submissions, as none of the target submissions are explicitly given to us. So we have to do some parsing of the target folder to get it all to work.","metadata":{}},{"cell_type":"code","source":"import json\n\nwith open(os.path.join(\"/kaggle/input/birdclef-2022\", \"scored_birds.json\")) as fp:\n    scored_birds = json.load(fp)\n\nprint(scored_birds)\n\n#submission_files = fastaudio.all.get_audio_files( '../input/birdclef-2022/test_soundscapes')\n\n#sub_dls = aud_digit.dataloaders(submission_files)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"submission = []\n\nfor i, row in tqdm(df_converted.iterrows(), total=len(df_converted)):\n    preds = dict(zip(model.labels, predictions[i]))\n    for bird in scored_birds:\n        submission.append({\n            \"row_id\": f\"{row['file_id']}_{bird}_{row['end_time']}\",\n            \"target\": preds[bird] > 1. / len(model.labels),\n        })","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"\nipd.Audio(\"../input/birdclef-2022/test_soundscapes/soundscape_453028782\")","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from pathlib import Path\nDATADIR = Path(\"../input/birdclef-2022/test_soundscapes/\")\nall_audios = list(DATADIR.glob(\"*.ogg\"))\nall_audios","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"submission_files","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_data","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"submission_sample","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# What to do now?\nWell, there are a couple of options:\n* Try tuning the image processing methods further\n* Doing more EDA to see if there are trends, be it in sound or in the metadata\n* Try seeing if there are other parameters we can use to optimize the learning more\n* Try alternate approaches <- Recommended idea\n\nHere are some notebooks that seemed cool, and focused mainly on alternate approaches:\n* [Focuses on using the lightning flash, a deep learning library](https://www.kaggle.com/code/jirkaborovec/birdclef-eda-baseline-flash-efficientnet/notebook)\n* These three all use PyTorch audio to some level:\n    * https://www.kaggle.com/code/tattaka/birdclef2022-submission-baseline\n    * https://www.kaggle.com/code/myso1987/pytorch-simple-starter-using-only-21-classes\n    * https://www.kaggle.com/code/shreyasajal/audio-albumentations-torchaudio-audiomentations\n\n","metadata":{}},{"cell_type":"code","source":"# An example of something you can do to train the model for longer\n# I kept them all at one step during the notebook so that it wouldn't take forever\n# You'll definitely not finish soon though!\n#learn.fit_one_cycle(10, sr.lr_steep)","metadata":{"trusted":true},"execution_count":null,"outputs":[]}]}