{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"> This notebook follows the [fastai style conventions](https://docs.fast.ai/dev/style.html#style-guide).","metadata":{}},{"cell_type":"markdown","source":"In this to-the-point notebook, I go over how one can create images of spectrograms from audio files using the PyTorch torchaudio module.\n\nThe notebook also goes over how I created the spectrogram images for the BirdCLEF 2023 competition, and how one can create and push a dataset right on Kaggle (useful if your local machine doesn't have enough storage).\n\nYou can view the dataset that was generated from this notebook [here](https://www.kaggle.com/datasets/forbo7/spectrograms-birdclef-2023).","metadata":{}},{"cell_type":"markdown","source":"## Setup","metadata":{}},{"cell_type":"code","source":"try: from fastkaggle import *\nexcept ModuleNotFoundError:\n    ! pip install -Uqq fastkaggle\n    from fastkaggle import *\n\niskaggle","metadata":{"execution":{"iopub.status.busy":"2023-04-04T08:47:01.961806Z","iopub.execute_input":"2023-04-04T08:47:01.962457Z","iopub.status.idle":"2023-04-04T08:47:16.995330Z","shell.execute_reply.started":"2023-04-04T08:47:01.962410Z","shell.execute_reply":"2023-04-04T08:47:16.993352Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"comp = 'birdclef-2023'\nd_path = setup_comp(comp, install='nbdev')","metadata":{"execution":{"iopub.status.busy":"2023-04-04T08:47:16.998131Z","iopub.execute_input":"2023-04-04T08:47:16.999269Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from fastai.imports import *\nfrom fastai.vision.all import *","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Data","metadata":{}},{"cell_type":"markdown","source":"### Paths","metadata":{}},{"cell_type":"markdown","source":"Let's see all the files and directories we have.","metadata":{}},{"cell_type":"code","source":"d_path.ls()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Let's get the path to the audio files.","metadata":{}},{"cell_type":"code","source":"aud_files = d_path/'train_audio'","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"And create a directory to store the spectrogram images.","metadata":{}},{"cell_type":"code","source":"mkdir('/kaggle/train_images', exist_ok=True); Path('/kaggle/train_images').exists()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Single Image","metadata":{}},{"cell_type":"markdown","source":"It's always a good idea to try things out on a smaller scale; so let's begin by converting only a single audio file.\n\nLet's get the first audio file.","metadata":{}},{"cell_type":"code","source":"aud_files.ls()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"aud_files.ls()[0].ls()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"aud = aud_files.ls()[0].ls()[0]; aud","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now it's time to load it in. What we get in return is the waveform and the sample rate.","metadata":{}},{"cell_type":"code","source":"import torchaudio\nwvfrm, sr = torchaudio.load(aud); wvfrm","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"---","metadata":{}},{"cell_type":"markdown","source":"#### BirdCLEF 2023 — Clipping the Audio Files","metadata":{}},{"cell_type":"markdown","source":"This competition requires predictions to be submitted of all 5 second intervals in each audio clip. This means the audio files need to be clipped.\n\nBelow is an easy way this can be done. We clip the first 5 seconds of the audio file.","metadata":{}},{"cell_type":"code","source":"start_sec = 0\nend_sec = 5\nwvfrm = wvfrm[:, start_sec*sr:end_sec*sr]\nwvfrm.shape[1] / sr","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Sample rate is simply the number of frames recorded per second. The waveform that torchaudio returns is a tensor of frames. Therefore, we can easily select the desired range of frames by multiplying the sample rate with the desired start and end seconds.","metadata":{}},{"cell_type":"markdown","source":"---","metadata":{}},{"cell_type":"markdown","source":"Now let's create the spectrogram.","metadata":{}},{"cell_type":"code","source":"import torchaudio.transforms as T\nspec = T.Spectrogram()(wvfrm); spec","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Let's scale it logarithmically. This allows for better viewing.","metadata":{}},{"cell_type":"code","source":"spec = T.AmplitudeToDB()(spec); spec","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The PyTorch tensor needs to be converted into a NumPy array so it can then further be converted to an image. I'm using the `squeeze` method to remove the uneeded axis of length 1, as seen below.","metadata":{}},{"cell_type":"code","source":"spec.shape","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"spec = spec.squeeze().numpy(); spec","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"spec.shape","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The array now needs to be normalized so it contains integers between 0 and 255: the values needed for images.","metadata":{}},{"cell_type":"code","source":"spec = (spec - spec.min()) / (spec.max() - spec.min()) * 255; spec","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"spec = spec.astype('uint8'); spec","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now we can finally convert the array to an image!","metadata":{}},{"cell_type":"code","source":"img = Image.fromarray(spec)\nprint(img.shape)\nimg","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Cool, hey? We've just visualized audio!","metadata":{}},{"cell_type":"markdown","source":"---","metadata":{}},{"cell_type":"markdown","source":"#### BirdCLEF 2023 — Resizing the Images","metadata":{}},{"cell_type":"markdown","source":"To allow the images to easily be used by various models, I resized the spectrograms to be 512 by 512 pixels as shown below.","metadata":{}},{"cell_type":"code","source":"img_size = (512, 512)\nimg = img.resize(img_size)\nprint(img.shape); img","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"---","metadata":{}},{"cell_type":"markdown","source":"To save the image, we can simply use the `save` method.","metadata":{}},{"cell_type":"code","source":"img.save('img.png')","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### All the Images","metadata":{}},{"cell_type":"markdown","source":"Now that we have verified that our algorithm works fine, we can extend it to convert all audio files.","metadata":{}},{"cell_type":"code","source":"def create_imgs(duration, f):\n    for step in range(0, duration, 5):\n        wvfrm, sr = torchaudio.load(f)\n        wvfrm = cut_wvfrm(wvfrm, sr, step)\n        spec = create_spec(wvfrm)\n        img = spec2img(spec)\n        end_sec = step + 5\n        img.save(f'/kaggle/train_images/{bird.stem}/{f.stem}_{end_sec}.png')\n\ndef cut_wvfrm(wvfrm, sr, step):\n    start_sec, end_sec = step, step + 5\n    return wvfrm[:, start_sec * sr: end_sec * sr]\n            \ndef create_spec(wvfrm):\n    spec = T.Spectrogram()(wvfrm)\n    return T.AmplitudeToDB()(spec)\n        \ndef spec2img(spec, img_size=(512, 512)):\n    spec = np.real(spec.squeeze().numpy())\n    spec = ((spec - spec.min()) / (spec.max() - spec.min()) * 255).astype('uint8')\n    return Image.fromarray(spec).resize(img_size)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"if not iskaggle:\n    for bird in aud_files.ls().sorted():\n        mkdir(f'/kaggle/train_images/{bird.stem}', exist_ok=True)\n        for f in bird.ls().sorted():\n            info = torchaudio.info(f)\n            duration = info.num_frames / info.sample_rate\n            if duration >= 5:\n                create_imgs(round(duration/5)*5, f)\n            else: continue","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"> **Note:** Ignore the `if not iskaggle` statement when replicating. I added it since I edited this notebook and needed to save changes without reproducing the entire dataset.","metadata":{}},{"cell_type":"markdown","source":"In the first `for` loop below, we loop through all the bird folders. For each folder, a folder with the same name is created in the directory where we want to store the images.\n\nIn the second `for` loop, we loop through all audio files within the folder and then convert them to spectrogram images through the `create_images` function I defined.","metadata":{}},{"cell_type":"markdown","source":"---","metadata":{}},{"cell_type":"markdown","source":"#### BirdCLEF 2023 — Clipping the Audio Files","metadata":{}},{"cell_type":"markdown","source":"Some audio files in the training set are of different durations. Therefore, we obtain the duration of the audio file so it can correctly be clipped into 5 second intervals.\n\n```python\ninfo = torchaudio.info(f)\nduration = info.num_frames / info.sample_rate\nif duration >= 5:\n    create_images(round(duration/5)*5, f)\nelse: continue\n```\n\nAgain, since sample rate is the number of frames recorded per second, we can divide the total number of frames by the sample rate to obtain the duration in seconds of a clip.\n\n`duration = info.num_frames / info.sample_rate`\n\nThen we round the duration to the nearest 5 for easy clipping.\n\n`round(duration/5)*5`","metadata":{}},{"cell_type":"markdown","source":"---","metadata":{}},{"cell_type":"markdown","source":"The images now created! The rest of this notebook covers how one can generate a dataset in a Kaggle Notebook and push it directly to Kaggle within it.","metadata":{}},{"cell_type":"markdown","source":"## API Setup","metadata":{}},{"cell_type":"markdown","source":"We need to configure the user keys so we can push to the correct account.\n\nTo do this, first obtain your Kaggle API key. Then, while in the notebook editor, click Add-ons -> Secrets -> Add a New Secret...\n\n<img src=\"https://sat02pap002files.storage.live.com/y4mY_6mu6XXxmpkR_MwTxHxfEGRV9QOqlxkCF8i6T90WkRadthJc9yeOxUG3k21bT-QQ4CxJcP1W-MHuAhX_tKvf-gcZqUvkLONiKPeumqZCTezKhwXJXs0OFuTDBnedbcViCOeeIkzjVjYJiBc3J7ib0AqG5v247vovXQP1JYv-8RmfWMFytuyCgwvLNt_5MSP?width=2880&height=1800&cropmode=none\" width=\"75%\" height=\"100%\" />\n\n<img src=\"https://sat02pap002files.storage.live.com/y4mBCKZg2xbmSClGOAWdxf6vXcrgLKV0TNe-SuLaS37nzKAuHpzMZsXL0lEAuauWYrrKnZVx2qGStH3ZjVn-E0xARzTUW7UN9Em3gx_X3qivCHIX_85Z0aQc0xQYrl2dJPvyPOHwuCm0ML5188bFuemgoQik3jn_TVih1YKxYw9Z79fDahzG8wD9Yu-zbizRsVS?width=1024&height=640&cropmode=none\" width=\"75%\" height=\"100%\" />\n\n<img src=\"https://sat02pap002files.storage.live.com/y4mgCDqODGoakfC1zbSK1tEGy1gcUPMpzzw5DEFHJnQPZIWZpWn2WKLT48EGRky-kXSmBYt-BYSgVxO7jMVT321_gqOYjX0daxumANws9LpNOvwdG3ce9nKJFT8ZkSEjtb5Q93leY75dKLk2CumNpmRGi4WoZr5jsTPPEV42yTlNPURAMrSsl09XeKqGaBjJO0u?width=2880&height=1800&cropmode=none\" width=\"75%\" height=\"100%\" />\n\n...input your key and give it a name...\n\n<img src=\"https://sat02pap002files.storage.live.com/y4mz3E2oKRG9H8n6cSwx-uk8ZAuIB8WSExJX92_w8ZoqYkX4f72QXKM2AFjjn7n_UaS30btHk6ueSi9VTRD_us_pUp3CSo3MkBM0mRBZqeqYDrSY4GRGamkrYnNKtxfZ-PQKJEvjOr8r8N97cftsJ3-9fEUcO3jv8kVGniKNPHWx_aVpIElx4_qgVq-aRt30PSD?width=2880&height=1800&cropmode=none\" width=\"75%\" height=\"100%\" />\n\n...and click save. Then click the checkbox next to the secret to activate it for your notebook.\n\n<img src=\"https://sat02pap002files.storage.live.com/y4m3RfkLD7cbynZIPmI7X17qmHTxar-hX5p-w5G70Z3YiRiIPzkzo3Z9h1Z4eWecrmmU88yl-HXUSRU7xUQ0LN2C-_wn_qmKDJFJyXSjNfhV7XjKpGbyhxXU9JCaTOxZhRwHGl4ghkykkhpMBTpoG-dxe7YlhPI40DQJT2dvCcuv1KZQkYkcOHIqiUI8SvwhL4Q?width=2880&height=1800&cropmode=none\" width=\"75%\" height=\"100%\" />\n\nRepeat for your Kaggle username.","metadata":{}},{"cell_type":"markdown","source":"Now we can set the keys for the notebook as shown below (input the name of your key into `get_secret`).","metadata":{}},{"cell_type":"code","source":"import os\nfrom kaggle_secrets import UserSecretsClient","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"secrets = UserSecretsClient()\nos.environ['KAGGLE_USERNAME'] = secrets.get_secret('KAGGLE_USERNAME')\nos.environ['KAGGLE_KEY'] = secrets.get_secret('KAGGLE_KEY')","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Push Dataset","metadata":{}},{"cell_type":"markdown","source":"The [fastkaggle](fastkaggle) library offers a convenient way to easily create and push a dataset to Kaggle.","metadata":{}},{"cell_type":"code","source":"doc(mk_dataset)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"> **Note:** Ignore the `if not iskaggle` statement when replicating. I added it since I edited this notebook and needed to save changes without reproducing the entire dataset.","metadata":{}},{"cell_type":"code","source":"if not iskaggle:\n    mk_dataset('/kaggle/train_images', 'spectrograms-birdclef-2023', force=True, upload=True)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"And we can verify our dataset has been created by having a look at the generated metadata file.","metadata":{}},{"cell_type":"code","source":"if not iskaggle:\n    ! cat /kaggle/train_images/dataset-metadata.json","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"From here, we can go directly to the dataset page on Kaggle and fill out the rest of the details.","metadata":{}},{"cell_type":"markdown","source":"## And there you have it!","metadata":{}},{"cell_type":"markdown","source":"In summary, you saw how to:\n* Generate spectrogram images from audio files using torchaudio and fastai\n* How to cut audio tracks\n* And how to create and push a dataset directly on Kaggle\n\nYou can view the dataset that was generated from this notebook [here](https://www.kaggle.com/datasets/forbo7/spectrograms-birdclef-2023).\n\nIf you have any comments, questions, suggestions, feedback, criticisms, or corrections, please do post them down in the comment section below!\n\nAnd if you found this helpful, please give this notebook an upvote so it can help others!","metadata":{}}]}