{"metadata":{"kaggle":{"accelerator":"nvidiaTeslaT4","dataSources":[{"sourceId":70203,"databundleVersionId":8068726,"sourceType":"competition"},{"sourceId":8084829,"sourceType":"datasetVersion","datasetId":4772325}],"dockerImageVersionId":30684,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":true},"kernelspec":{"display_name":"Python 3 (ipykernel)","language":"python","name":"python3"},"language_info":{"codemirror_mode":{"name":"ipython","version":3},"file_extension":".py","mimetype":"text/x-python","name":"python","nbconvert_exporter":"python","pygments_lexer":"ipython3","version":"3.9.13"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"code","source":"# kaggle installs\n# !pip install /kaggle/input/fastxtend-0-1-7/fastxtend-0.1.7-py3-none-any.whl -qq\n# !pip install /kaggle/input/fastxtend-0-1-7/colorednoise-2.2.0-py3-none-any.whl -qq\n# !pip install /kaggle/input/fastxtend-0-1-7/primePy-1.3-py3-none-any.whl -qq","metadata":{"_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","execution":{"iopub.execute_input":"2024-04-15T15:34:34.770743Z","iopub.status.busy":"2024-04-15T15:34:34.769906Z","iopub.status.idle":"2024-04-15T15:35:13.708551Z","shell.execute_reply":"2024-04-15T15:35:13.707197Z","shell.execute_reply.started":"2024-04-15T15:34:34.770708Z"},"_kg_hide-input":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#!zip -r /notebooks/train_chunks.zip /notebooks/train_chunks","metadata":{"execution":{"iopub.execute_input":"2024-04-16T14:49:08.373884Z","iopub.status.busy":"2024-04-16T14:49:08.373254Z","iopub.status.idle":"2024-04-16T14:49:08.377619Z","shell.execute_reply":"2024-04-16T14:49:08.376778Z","shell.execute_reply.started":"2024-04-16T14:49:08.373852Z"},"_kg_hide-input":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# paperspace installs\n# !pip install fastai==2.7.14 torch==2.1.2 torchaudio==2.1.2 -qq\n# !pip install /notebooks/wheels/primePy-1.3-py3-none-any.whl -qq\n# !pip install /notebooks/wheels/colorednoise-2.2.0-py3-none-any.whl -qq\n# !pip install /notebooks/wheels/fastxtend-0.1.7-py3-none-any.whl -qq","metadata":{"execution":{"iopub.execute_input":"2024-04-15T19:37:37.031475Z","iopub.status.busy":"2024-04-15T19:37:37.030939Z","iopub.status.idle":"2024-04-15T19:37:44.190219Z","shell.execute_reply":"2024-04-15T19:37:44.189194Z","shell.execute_reply.started":"2024-04-15T19:37:37.031452Z"},"_kg_hide-input":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from fastai.vision.all import *\nfrom fastxtend.audio.all import *\nfrom fastcore.parallel import *\nimport librosa\nimport ast","metadata":{"execution":{"iopub.execute_input":"2024-04-15T19:47:09.067251Z","iopub.status.busy":"2024-04-15T19:47:09.066975Z","iopub.status.idle":"2024-04-15T19:47:09.072328Z","shell.execute_reply":"2024-04-15T19:47:09.071562Z","shell.execute_reply.started":"2024-04-15T19:47:09.067231Z"},"_kg_hide-input":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#!kaggle competitions download -c birdclef-2024\n#import zipfile\n#zipfile.ZipFile('birdclef-2024.zip').extractall('birdclef-2024')","metadata":{"execution":{"iopub.execute_input":"2024-04-16T14:49:21.016627Z","iopub.status.busy":"2024-04-16T14:49:21.016360Z","iopub.status.idle":"2024-04-16T14:49:21.019901Z","shell.execute_reply":"2024-04-16T14:49:21.019290Z","shell.execute_reply.started":"2024-04-16T14:49:21.016609Z"},"_kg_hide-input":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# kaggle path\n#path = Path('/kaggle/input/birdclef-2024')\n\n# paperspace path\npath = Path('/notebooks/birdclef-2024')\nfor el in path.ls():\n    print(el)","metadata":{"execution":{"iopub.execute_input":"2024-04-15T19:47:30.880777Z","iopub.status.busy":"2024-04-15T19:47:30.880204Z","iopub.status.idle":"2024-04-15T19:47:30.887695Z","shell.execute_reply":"2024-04-15T19:47:30.887135Z","shell.execute_reply.started":"2024-04-15T19:47:30.880755Z"},"_kg_hide-input":true,"_kg_hide-output":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Background","metadata":{}},{"cell_type":"markdown","source":"In my [previous notebook](https://www.kaggle.com/code/vishalbakshi/birdclef2024-getting-started-with-fastxtend-audio) I got familiar with using fastxtend to train an image classifier on audio spectrograms. In this notebook I'll train a quick-and-dirty model on BirdCLEF 2024 competition data and export it so I can use it submit predictions in the next notebook (with internet access disabled). Note that this notebook was created and run mostly in Paperspace so the filepaths (and training times) are all based on that workspace.","metadata":{}},{"cell_type":"markdown","source":"## Understand the Dataset","metadata":{}},{"cell_type":"markdown","source":"I'll understand the competition dataset more and more as I train different models over the next two months, but would like to get some high level features about the training dataset understood right now:\n\n- How many audio files are there?\n- How many classes?\n- How many audio files per class?\n- How many types of sound?\n- How many authors?\n- How many files per rating?\n- How many secondary labels?\n- How long are the audio files?\n- What is the sample rate of the audio files?","metadata":{}},{"cell_type":"code","source":"df = pd.read_csv(path/'train_metadata.csv')\nprint('Number of audio files:', len(df))\ndf.head()","metadata":{"execution":{"iopub.execute_input":"2024-04-15T19:49:13.997020Z","iopub.status.busy":"2024-04-15T19:49:13.996115Z","iopub.status.idle":"2024-04-15T19:49:14.131884Z","shell.execute_reply":"2024-04-15T19:49:14.131178Z","shell.execute_reply.started":"2024-04-15T19:49:13.996993Z"}},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print('Number of classes:', len(df.primary_label.unique()))","metadata":{"execution":{"iopub.execute_input":"2024-04-15T19:49:15.016611Z","iopub.status.busy":"2024-04-15T19:49:15.015957Z","iopub.status.idle":"2024-04-15T19:49:15.024662Z","shell.execute_reply":"2024-04-15T19:49:15.023787Z","shell.execute_reply.started":"2024-04-15T19:49:15.016587Z"}},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(\"Number of audio files per class\")\nprint(df.groupby('primary_label')['primary_label'].count().median())","metadata":{"execution":{"iopub.execute_input":"2024-04-15T19:49:15.371935Z","iopub.status.busy":"2024-04-15T19:49:15.371012Z","iopub.status.idle":"2024-04-15T19:49:15.380468Z","shell.execute_reply":"2024-04-15T19:49:15.379795Z","shell.execute_reply.started":"2024-04-15T19:49:15.371908Z"}},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(df.groupby('primary_label')['primary_label'].count().hist())","metadata":{"execution":{"iopub.execute_input":"2024-04-15T19:49:15.649034Z","iopub.status.busy":"2024-04-15T19:49:15.648108Z","iopub.status.idle":"2024-04-15T19:49:15.791539Z","shell.execute_reply":"2024-04-15T19:49:15.790825Z","shell.execute_reply.started":"2024-04-15T19:49:15.648998Z"}},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(df.groupby('primary_label')['primary_label'].count().min(), df.groupby('primary_label')['primary_label'].count().max())","metadata":{"execution":{"iopub.execute_input":"2024-04-15T19:49:16.129470Z","iopub.status.busy":"2024-04-15T19:49:16.129192Z","iopub.status.idle":"2024-04-15T19:49:16.137990Z","shell.execute_reply":"2024-04-15T19:49:16.137472Z","shell.execute_reply.started":"2024-04-15T19:49:16.129452Z"}},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The `type` column consists of strings of lists:","metadata":{}},{"cell_type":"code","source":"df.type.iloc[7]","metadata":{"execution":{"iopub.execute_input":"2024-04-15T19:49:19.114631Z","iopub.status.busy":"2024-04-15T19:49:19.113753Z","iopub.status.idle":"2024-04-15T19:49:19.119543Z","shell.execute_reply":"2024-04-15T19:49:19.118347Z","shell.execute_reply.started":"2024-04-15T19:49:19.114585Z"}},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"I'll convert those strings to actual `list`s first:","metadata":{}},{"cell_type":"code","source":"type_vals = [ast.literal_eval(val) for val in df.type.unique()]\ntype_vals[:5]","metadata":{"execution":{"iopub.execute_input":"2024-04-15T19:49:19.816322Z","iopub.status.busy":"2024-04-15T19:49:19.815647Z","iopub.status.idle":"2024-04-15T19:49:19.829324Z","shell.execute_reply":"2024-04-15T19:49:19.828783Z","shell.execute_reply.started":"2024-04-15T19:49:19.816296Z"}},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Then, I'll create a flattened list of all string values:","metadata":{}},{"cell_type":"code","source":"flattened_types = [item for sublist in type_vals for item in sublist]\nflattened_types[:5]","metadata":{"execution":{"iopub.execute_input":"2024-04-15T19:49:20.536034Z","iopub.status.busy":"2024-04-15T19:49:20.535379Z","iopub.status.idle":"2024-04-15T19:49:20.540567Z","shell.execute_reply":"2024-04-15T19:49:20.539958Z","shell.execute_reply.started":"2024-04-15T19:49:20.536012Z"}},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Finally, I'll count the number of unique `type` values:","metadata":{}},{"cell_type":"code","source":"len(flattened_types)","metadata":{"execution":{"iopub.execute_input":"2024-04-15T19:49:21.012495Z","iopub.status.busy":"2024-04-15T19:49:21.011822Z","iopub.status.idle":"2024-04-15T19:49:21.016345Z","shell.execute_reply":"2024-04-15T19:49:21.015864Z","shell.execute_reply.started":"2024-04-15T19:49:21.012474Z"}},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print('Number of sound types:',  len(pd.Series(flattened_types).unique()))","metadata":{"execution":{"iopub.execute_input":"2024-04-15T19:49:21.276650Z","iopub.status.busy":"2024-04-15T19:49:21.276387Z","iopub.status.idle":"2024-04-15T19:49:21.281298Z","shell.execute_reply":"2024-04-15T19:49:21.280520Z","shell.execute_reply.started":"2024-04-15T19:49:21.276632Z"}},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print('Number of authors:', len(df.author.unique()))","metadata":{"execution":{"iopub.execute_input":"2024-04-15T19:49:21.501298Z","iopub.status.busy":"2024-04-15T19:49:21.501039Z","iopub.status.idle":"2024-04-15T19:49:21.506026Z","shell.execute_reply":"2024-04-15T19:49:21.505485Z","shell.execute_reply.started":"2024-04-15T19:49:21.501280Z"}},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print('Mean and Median Rating:', df.rating.mean(), ',', df.rating.median())","metadata":{"execution":{"iopub.execute_input":"2024-04-15T19:49:21.578965Z","iopub.status.busy":"2024-04-15T19:49:21.578354Z","iopub.status.idle":"2024-04-15T19:49:21.583203Z","shell.execute_reply":"2024-04-15T19:49:21.582716Z","shell.execute_reply.started":"2024-04-15T19:49:21.578944Z"}},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print('Number of files per rating:', df.rating.hist())","metadata":{"execution":{"iopub.execute_input":"2024-04-15T19:49:21.753734Z","iopub.status.busy":"2024-04-15T19:49:21.753092Z","iopub.status.idle":"2024-04-15T19:49:21.824391Z","shell.execute_reply":"2024-04-15T19:49:21.823885Z","shell.execute_reply.started":"2024-04-15T19:49:21.753710Z"}},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The majority of ratings are greater than or equal to 3.5:","metadata":{}},{"cell_type":"code","source":"len(df.query('rating >=3.5'))/len(df)","metadata":{"execution":{"iopub.execute_input":"2024-04-15T19:49:25.919433Z","iopub.status.busy":"2024-04-15T19:49:25.918894Z","iopub.status.idle":"2024-04-15T19:49:25.931947Z","shell.execute_reply":"2024-04-15T19:49:25.930809Z","shell.execute_reply.started":"2024-04-15T19:49:25.919410Z"}},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"A relatively small percentage of audio files are multi-category. Something to explore later on.","metadata":{}},{"cell_type":"code","source":"len(df.query(\"secondary_labels != '[]'\"))/len(df)","metadata":{"execution":{"iopub.execute_input":"2024-04-15T19:49:26.496654Z","iopub.status.busy":"2024-04-15T19:49:26.496379Z","iopub.status.idle":"2024-04-15T19:49:26.506371Z","shell.execute_reply":"2024-04-15T19:49:26.505803Z","shell.execute_reply.started":"2024-04-15T19:49:26.496633Z"}},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now I'll look at what I'm most interested in at this stage---how long are these audio files, and what is their sampling rate?","metadata":{}},{"cell_type":"code","source":"(path/'train_audio').ls()","metadata":{"execution":{"iopub.execute_input":"2024-04-15T19:49:27.043557Z","iopub.status.busy":"2024-04-15T19:49:27.042878Z","iopub.status.idle":"2024-04-15T19:49:27.051973Z","shell.execute_reply":"2024-04-15T19:49:27.051367Z","shell.execute_reply.started":"2024-04-15T19:49:27.043532Z"}},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fnames = get_files(path/'train_audio')","metadata":{"execution":{"iopub.execute_input":"2024-04-15T19:49:27.371780Z","iopub.status.busy":"2024-04-15T19:49:27.371105Z","iopub.status.idle":"2024-04-15T19:49:27.774232Z","shell.execute_reply":"2024-04-15T19:49:27.773574Z","shell.execute_reply.started":"2024-04-15T19:49:27.371725Z"}},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print('Number of files matches metadata length:', len(fnames) == len(df), '\\n', 'Number of audio files:', len(fnames))","metadata":{"execution":{"iopub.execute_input":"2024-04-15T19:49:27.776191Z","iopub.status.busy":"2024-04-15T19:49:27.775549Z","iopub.status.idle":"2024-04-15T19:49:27.779760Z","shell.execute_reply":"2024-04-15T19:49:27.779073Z","shell.execute_reply.started":"2024-04-15T19:49:27.776166Z"}},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"I'll use fastxtend's `TensorAudio` to capture the sampling rate and duration of the audio files. [Peeking under the hood](https://github.com/warner-benjamin/fastxtend/blob/ff594c81e16b1ca062744ab46bd519968d6ef4dc/fastxtend/audio/core.py#L50), `TensorAudio` instantiates by calling `torchaudio.load` and capturing the audio data and sampling rate. It calcuates duration by dividing the total number of samples by the sampling rate. ","metadata":{}},{"cell_type":"code","source":"ta = TensorAudio.create(fnames[0])\nta","metadata":{"execution":{"iopub.execute_input":"2024-04-15T19:49:30.646166Z","iopub.status.busy":"2024-04-15T19:49:30.645610Z","iopub.status.idle":"2024-04-15T19:49:30.713892Z","shell.execute_reply":"2024-04-15T19:49:30.713348Z","shell.execute_reply.started":"2024-04-15T19:49:30.646140Z"}},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print('Sample Rate:', ta.sr)","metadata":{"execution":{"iopub.execute_input":"2024-04-15T19:49:30.997834Z","iopub.status.busy":"2024-04-15T19:49:30.997205Z","iopub.status.idle":"2024-04-15T19:49:31.000912Z","shell.execute_reply":"2024-04-15T19:49:31.000349Z","shell.execute_reply.started":"2024-04-15T19:49:30.997812Z"}},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print('Duration (seconds):', ta.duration)","metadata":{"execution":{"iopub.execute_input":"2024-04-15T19:49:31.229720Z","iopub.status.busy":"2024-04-15T19:49:31.229145Z","iopub.status.idle":"2024-04-15T19:49:31.233867Z","shell.execute_reply":"2024-04-15T19:49:31.232949Z","shell.execute_reply.started":"2024-04-15T19:49:31.229694Z"}},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"ta.show();","metadata":{"execution":{"iopub.execute_input":"2024-04-15T19:49:31.468528Z","iopub.status.busy":"2024-04-15T19:49:31.467878Z","iopub.status.idle":"2024-04-15T19:49:35.719163Z","shell.execute_reply":"2024-04-15T19:49:35.718563Z","shell.execute_reply.started":"2024-04-15T19:49:31.468506Z"}},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def f(o):\n    ta = TensorAudio.create(o)\n    return ta.sr, ta.duration","metadata":{"execution":{"iopub.execute_input":"2024-04-15T19:49:35.720789Z","iopub.status.busy":"2024-04-15T19:49:35.720512Z","iopub.status.idle":"2024-04-15T19:49:35.725002Z","shell.execute_reply":"2024-04-15T19:49:35.724080Z","shell.execute_reply.started":"2024-04-15T19:49:35.720768Z"}},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"f(fnames[0])","metadata":{"execution":{"iopub.execute_input":"2024-04-15T19:49:35.726645Z","iopub.status.busy":"2024-04-15T19:49:35.726113Z","iopub.status.idle":"2024-04-15T19:49:35.744083Z","shell.execute_reply":"2024-04-15T19:49:35.743576Z","shell.execute_reply.started":"2024-04-15T19:49:35.726619Z"}},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"audio_data = parallel(f, fnames, n_workers=4)","metadata":{"execution":{"iopub.execute_input":"2024-04-15T19:49:35.745794Z","iopub.status.busy":"2024-04-15T19:49:35.745300Z","iopub.status.idle":"2024-04-15T19:53:16.151396Z","shell.execute_reply":"2024-04-15T19:53:16.150586Z","shell.execute_reply.started":"2024-04-15T19:49:35.745773Z"}},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"audio_data_df = pd.DataFrame({'data': pd.Series(audio_data)})\naudio_data_df['sr'] = audio_data_df.data.apply(lambda x: x[0])\naudio_data_df['duration'] = audio_data_df.data.apply(lambda x: x[1])","metadata":{"execution":{"iopub.execute_input":"2024-04-15T19:53:16.152746Z","iopub.status.busy":"2024-04-15T19:53:16.152358Z","iopub.status.idle":"2024-04-15T19:53:16.174404Z","shell.execute_reply":"2024-04-15T19:53:16.173724Z","shell.execute_reply.started":"2024-04-15T19:53:16.152723Z"}},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"len(audio_data_df)","metadata":{"execution":{"iopub.execute_input":"2024-04-15T19:53:16.175617Z","iopub.status.busy":"2024-04-15T19:53:16.175380Z","iopub.status.idle":"2024-04-15T19:53:16.180291Z","shell.execute_reply":"2024-04-15T19:53:16.179732Z","shell.execute_reply.started":"2024-04-15T19:53:16.175593Z"}},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"audio_data_df.groupby('sr')['sr'].count()","metadata":{"execution":{"iopub.execute_input":"2024-04-15T19:53:16.181629Z","iopub.status.busy":"2024-04-15T19:53:16.181380Z","iopub.status.idle":"2024-04-15T19:53:16.188995Z","shell.execute_reply":"2024-04-15T19:53:16.188412Z","shell.execute_reply.started":"2024-04-15T19:53:16.181607Z"}},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"All of the audio file are sampled at 32 kHz.","metadata":{}},{"cell_type":"code","source":"audio_data_df.duration.hist();","metadata":{"execution":{"iopub.execute_input":"2024-04-15T19:53:16.190314Z","iopub.status.busy":"2024-04-15T19:53:16.190091Z","iopub.status.idle":"2024-04-15T19:53:16.348041Z","shell.execute_reply":"2024-04-15T19:53:16.347434Z","shell.execute_reply.started":"2024-04-15T19:53:16.190293Z"}},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Most of the audio files are 500 seconds of less--that's still a very wide range.","metadata":{}},{"cell_type":"code","source":"audio_data_df.query(\"duration <= 500\").duration.hist();","metadata":{"execution":{"iopub.execute_input":"2024-04-15T19:53:16.349282Z","iopub.status.busy":"2024-04-15T19:53:16.349087Z","iopub.status.idle":"2024-04-15T19:53:16.429687Z","shell.execute_reply":"2024-04-15T19:53:16.429154Z","shell.execute_reply.started":"2024-04-15T19:53:16.349266Z"}},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Most of the audio files are less than 100 seconds long:","metadata":{}},{"cell_type":"code","source":"len(audio_data_df.query(\"duration <= 100\"))/len(audio_data_df)","metadata":{"execution":{"iopub.execute_input":"2024-04-15T19:53:16.432074Z","iopub.status.busy":"2024-04-15T19:53:16.431862Z","iopub.status.idle":"2024-04-15T19:53:16.440681Z","shell.execute_reply":"2024-04-15T19:53:16.439822Z","shell.execute_reply.started":"2024-04-15T19:53:16.432057Z"}},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"audio_data_df.query(\"duration <= 100\").duration.hist();","metadata":{"execution":{"iopub.execute_input":"2024-04-15T19:53:16.442003Z","iopub.status.busy":"2024-04-15T19:53:16.441806Z","iopub.status.idle":"2024-04-15T19:53:16.518614Z","shell.execute_reply":"2024-04-15T19:53:16.517788Z","shell.execute_reply.started":"2024-04-15T19:53:16.441987Z"}},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"About half of the files are 25 seconds or shorter:","metadata":{}},{"cell_type":"code","source":"len(audio_data_df.query(\"duration <= 25\"))/len(df)","metadata":{"execution":{"iopub.execute_input":"2024-04-15T19:53:16.520516Z","iopub.status.busy":"2024-04-15T19:53:16.520235Z","iopub.status.idle":"2024-04-15T19:53:16.529650Z","shell.execute_reply":"2024-04-15T19:53:16.528869Z","shell.execute_reply.started":"2024-04-15T19:53:16.520491Z"}},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Lastly, I'll look at some summary statistics for the `duration` of audio files in the dataset:","metadata":{}},{"cell_type":"code","source":"audio_data_df.duration.describe()","metadata":{"execution":{"iopub.execute_input":"2024-04-15T19:53:16.531218Z","iopub.status.busy":"2024-04-15T19:53:16.530953Z","iopub.status.idle":"2024-04-15T19:53:16.540625Z","shell.execute_reply":"2024-04-15T19:53:16.539866Z","shell.execute_reply.started":"2024-04-15T19:53:16.531194Z"}},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"One initial thought I had was to split the audio files into 5-second segments to train on. There would be about 205k total such segments.","metadata":{}},{"cell_type":"code","source":"print('Number of 5-second segments:', audio_data_df.duration.sum()/5)","metadata":{"execution":{"iopub.execute_input":"2024-04-15T19:53:16.542092Z","iopub.status.busy":"2024-04-15T19:53:16.541817Z","iopub.status.idle":"2024-04-15T19:53:16.547042Z","shell.execute_reply":"2024-04-15T19:53:16.546329Z","shell.execute_reply.started":"2024-04-15T19:53:16.542068Z"}},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Using `RandomCropPad`","metadata":{}},{"cell_type":"markdown","source":"fastxtend has a `RandomCropPad` augmentation which will \"randomly resize `TensorAudio` to specified length\". I'll use it on the first file (which is about 11 seconds long):","metadata":{}},{"cell_type":"code","source":"ta.show();","metadata":{"execution":{"iopub.execute_input":"2024-04-15T19:53:16.548519Z","iopub.status.busy":"2024-04-15T19:53:16.548262Z","iopub.status.idle":"2024-04-15T19:53:16.774211Z","shell.execute_reply":"2024-04-15T19:53:16.773385Z","shell.execute_reply.started":"2024-04-15T19:53:16.548494Z"}},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"As done in the [docs](https://fastxtend.benjaminwarner.dev/audio.03_augment.html#randomcroppad), I'll see how each of the four `padmodes` look like for this crop of `5` seconds:","metadata":{}},{"cell_type":"code","source":"with less_random():\n    _,axs = plt.subplots(1,4,figsize=(18,4))\n    for ax,padmode in zip(axs.flatten(), [AudioPadMode.Constant, AudioPadMode.ConstantPre,\n                                          AudioPadMode.ConstantPost, AudioPadMode.Repeat]\n    ):\n        rcp = RandomCropPad(5, padmode=padmode)\n        rcp(ta, split_idx=1).show(ctx=ax, title=padmode, hear=False)\n        plt.tight_layout()","metadata":{"execution":{"iopub.execute_input":"2024-04-15T19:53:16.775904Z","iopub.status.busy":"2024-04-15T19:53:16.775586Z","iopub.status.idle":"2024-04-15T19:53:18.175897Z","shell.execute_reply":"2024-04-15T19:53:18.175361Z","shell.execute_reply.started":"2024-04-15T19:53:16.775877Z"}},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Since the audio file is longer than 5 seconds, the `padmodes` result in the same output. I'll look at a file that is fewer than 5 seconds:","metadata":{}},{"cell_type":"code","source":"audio_data_df.query(\"duration < 3\").iloc[0]","metadata":{"execution":{"iopub.execute_input":"2024-04-15T19:56:25.904642Z","iopub.status.busy":"2024-04-15T19:56:25.904010Z","iopub.status.idle":"2024-04-15T19:56:25.912464Z","shell.execute_reply":"2024-04-15T19:56:25.911949Z","shell.execute_reply.started":"2024-04-15T19:56:25.904618Z"}},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"f(fnames[22])","metadata":{"execution":{"iopub.execute_input":"2024-04-15T19:56:43.032291Z","iopub.status.busy":"2024-04-15T19:56:43.031521Z","iopub.status.idle":"2024-04-15T19:56:43.040678Z","shell.execute_reply":"2024-04-15T19:56:43.040103Z","shell.execute_reply.started":"2024-04-15T19:56:43.032266Z"}},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"with less_random():\n    _,axs = plt.subplots(1,4,figsize=(18,4))\n    for ax,padmode in zip(axs.flatten(), [AudioPadMode.Constant, AudioPadMode.ConstantPre,\n                                          AudioPadMode.ConstantPost, AudioPadMode.Repeat]\n    ):\n        rcp = RandomCropPad(5, padmode=padmode)\n        rcp(TensorAudio.create(fnames[22]), split_idx=1).show(ctx=ax, title=padmode, hear=False)\n        plt.tight_layout()","metadata":{"execution":{"iopub.execute_input":"2024-04-15T19:56:43.943603Z","iopub.status.busy":"2024-04-15T19:56:43.943085Z","iopub.status.idle":"2024-04-15T19:56:45.275260Z","shell.execute_reply":"2024-04-15T19:56:45.274811Z","shell.execute_reply.started":"2024-04-15T19:56:43.943578Z"}},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Here I can see more clearly how the different `padmode` types affect the audio signal.","metadata":{}},{"cell_type":"markdown","source":"## Preparing 5-second Audio Tensors","metadata":{}},{"cell_type":"markdown","source":"During inference, the test audio files provided are 4 minutes long and need to be cut up into 5-second chunks, upon which predictions will be calculated. The input to the model (and `DataLoader`) will therefore be `TensoAudio` objects.","metadata":{}},{"cell_type":"markdown","source":"I'll train on only a small subset (about 2000 or 10%) of the dataset so I can quickly train my initial model. I'll make sure to sample from each `primary_label` class so my validation calculations contain classes that are included during training.","metadata":{}},{"cell_type":"code","source":"grouped = df.groupby('primary_label')","metadata":{"execution":{"iopub.execute_input":"2024-04-15T19:56:49.536997Z","iopub.status.busy":"2024-04-15T19:56:49.536507Z","iopub.status.idle":"2024-04-15T19:56:49.539949Z","shell.execute_reply":"2024-04-15T19:56:49.539514Z","shell.execute_reply.started":"2024-04-15T19:56:49.536976Z"}},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def sample_from_group(group, n):\n    return group.sample(min(n, len(group)))","metadata":{"execution":{"iopub.execute_input":"2024-04-15T19:56:50.320593Z","iopub.status.busy":"2024-04-15T19:56:50.319786Z","iopub.status.idle":"2024-04-15T19:56:50.323267Z","shell.execute_reply":"2024-04-15T19:56:50.322851Z","shell.execute_reply.started":"2024-04-15T19:56:50.320568Z"}},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Number of rows to sample from each group\nsample_size_per_group = 10\n\n# Apply the sampling function to each group and concatenate the results\nsampled_df = grouped.apply(lambda group: sample_from_group(group, sample_size_per_group))\n\n# Reset the index if needed\nsampled_df.reset_index(drop=True, inplace=True)","metadata":{"execution":{"iopub.execute_input":"2024-04-15T19:56:51.146695Z","iopub.status.busy":"2024-04-15T19:56:51.145909Z","iopub.status.idle":"2024-04-15T19:56:51.184675Z","shell.execute_reply":"2024-04-15T19:56:51.184238Z","shell.execute_reply.started":"2024-04-15T19:56:51.146670Z"}},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sampled_df.head(3)","metadata":{"execution":{"iopub.execute_input":"2024-04-15T19:56:51.560547Z","iopub.status.busy":"2024-04-15T19:56:51.559832Z","iopub.status.idle":"2024-04-15T19:56:51.569150Z","shell.execute_reply":"2024-04-15T19:56:51.568556Z","shell.execute_reply.started":"2024-04-15T19:56:51.560523Z"}},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"len(sampled_df)","metadata":{"execution":{"iopub.execute_input":"2024-04-15T19:56:54.741809Z","iopub.status.busy":"2024-04-15T19:56:54.741262Z","iopub.status.idle":"2024-04-15T19:56:54.746270Z","shell.execute_reply":"2024-04-15T19:56:54.745731Z","shell.execute_reply.started":"2024-04-15T19:56:54.741788Z"}},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"len(sampled_df.primary_label.unique())","metadata":{"execution":{"iopub.execute_input":"2024-04-15T19:56:57.603839Z","iopub.status.busy":"2024-04-15T19:56:57.603562Z","iopub.status.idle":"2024-04-15T19:56:57.608748Z","shell.execute_reply":"2024-04-15T19:56:57.608028Z","shell.execute_reply.started":"2024-04-15T19:56:57.603820Z"}},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Great---my sampled `DataFrame` contains all 182 classes.","metadata":{}},{"cell_type":"markdown","source":"### Creating Chunk Filenames","metadata":{}},{"cell_type":"markdown","source":"Next, I'll prepare a `Series` of filenames that are in the format `{class}/{stem}_{end time}.pt`, where `end time` is a multiple of 5. I'll start by adding a column `n_chunks` to `sampled_df`, which contains the integer of 5 second chunks.","metadata":{}},{"cell_type":"code","source":"sampled_df['n_chunks'] = sampled_df.filename.apply(lambda x: int(TensorAudio.create(path/'train_audio'/x).duration) // 5)","metadata":{"execution":{"iopub.execute_input":"2024-04-15T19:56:59.199014Z","iopub.status.busy":"2024-04-15T19:56:59.198752Z","iopub.status.idle":"2024-04-15T19:57:43.955384Z","shell.execute_reply":"2024-04-15T19:57:43.954743Z","shell.execute_reply.started":"2024-04-15T19:56:59.198995Z"}},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The first audio file is 38 seconds long so the `n_chunks` value should be 7.","metadata":{}},{"cell_type":"code","source":"path/'train_audio'/sampled_df.filename.iloc[0]","metadata":{"execution":{"iopub.execute_input":"2024-04-15T19:57:43.959283Z","iopub.status.busy":"2024-04-15T19:57:43.958800Z","iopub.status.idle":"2024-04-15T19:57:43.966212Z","shell.execute_reply":"2024-04-15T19:57:43.965693Z","shell.execute_reply.started":"2024-04-15T19:57:43.959263Z"}},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"int(TensorAudio.create(path/'train_audio'/sampled_df.filename.iloc[0]).duration) // 5","metadata":{"execution":{"iopub.execute_input":"2024-04-15T19:57:43.967087Z","iopub.status.busy":"2024-04-15T19:57:43.966914Z","iopub.status.idle":"2024-04-15T19:57:44.004189Z","shell.execute_reply":"2024-04-15T19:57:44.003698Z","shell.execute_reply.started":"2024-04-15T19:57:43.967072Z"}},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sampled_df.head(1)","metadata":{"execution":{"iopub.execute_input":"2024-04-15T19:57:44.005758Z","iopub.status.busy":"2024-04-15T19:57:44.005566Z","iopub.status.idle":"2024-04-15T19:57:44.014614Z","shell.execute_reply":"2024-04-15T19:57:44.014030Z","shell.execute_reply.started":"2024-04-15T19:57:44.005742Z"}},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Looks good, I'll now create the `Series` of chunk filenames for all audio files:","metadata":{}},{"cell_type":"code","source":"chunk_fnames = pd.Series([f\"{row['filename'].split('/')[1].split('.')[0]}_{end_time}\" \n                          for _, row in sampled_df.iterrows() \n                          for end_time in [(i + 1) * 5 \n                                           for i in range(row['n_chunks'])]])","metadata":{"execution":{"iopub.execute_input":"2024-04-15T19:57:44.015778Z","iopub.status.busy":"2024-04-15T19:57:44.015602Z","iopub.status.idle":"2024-04-15T19:57:44.109973Z","shell.execute_reply":"2024-04-15T19:57:44.109358Z","shell.execute_reply.started":"2024-04-15T19:57:44.015763Z"}},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The first audio file is 38 seconds long so the last 5-second chunk should end with `_35`.","metadata":{}},{"cell_type":"code","source":"chunk_fnames[:8]","metadata":{"execution":{"iopub.execute_input":"2024-04-15T19:57:44.110851Z","iopub.status.busy":"2024-04-15T19:57:44.110678Z","iopub.status.idle":"2024-04-15T19:57:44.123541Z","shell.execute_reply":"2024-04-15T19:57:44.123066Z","shell.execute_reply.started":"2024-04-15T19:57:44.110835Z"}},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The last audio file as 4 chunks:","metadata":{}},{"cell_type":"code","source":"sampled_df.iloc[-1]","metadata":{"execution":{"iopub.execute_input":"2024-04-15T19:57:44.124479Z","iopub.status.busy":"2024-04-15T19:57:44.124273Z","iopub.status.idle":"2024-04-15T19:57:44.129511Z","shell.execute_reply":"2024-04-15T19:57:44.129021Z","shell.execute_reply.started":"2024-04-15T19:57:44.124461Z"}},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"chunk_fnames[-5:]","metadata":{"execution":{"iopub.execute_input":"2024-04-15T19:57:44.130302Z","iopub.status.busy":"2024-04-15T19:57:44.130146Z","iopub.status.idle":"2024-04-15T19:57:44.135327Z","shell.execute_reply":"2024-04-15T19:57:44.134870Z","shell.execute_reply.started":"2024-04-15T19:57:44.130288Z"}},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The number of `chunk_fnames` should equal the `sum` of `n_chunks`:","metadata":{}},{"cell_type":"code","source":"len(chunk_fnames) == sampled_df.n_chunks.sum()","metadata":{"execution":{"iopub.execute_input":"2024-04-15T19:57:44.136553Z","iopub.status.busy":"2024-04-15T19:57:44.136391Z","iopub.status.idle":"2024-04-15T19:57:44.140407Z","shell.execute_reply":"2024-04-15T19:57:44.139905Z","shell.execute_reply.started":"2024-04-15T19:57:44.136538Z"}},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Creating Chunk Labels","metadata":{}},{"cell_type":"markdown","source":"Next, I'll prepare the `primary_label`s for the chunks (which will be the `primary_label` column repeated for the number of chunks for each audio file).","metadata":{}},{"cell_type":"code","source":"chunk_labels = pd.Series([row['primary_label']\n                         for _, row in sampled_df.iterrows()\n                         for chunk in range(row['n_chunks'])])","metadata":{"execution":{"iopub.execute_input":"2024-04-15T19:57:44.142414Z","iopub.status.busy":"2024-04-15T19:57:44.142063Z","iopub.status.idle":"2024-04-15T19:57:44.233313Z","shell.execute_reply":"2024-04-15T19:57:44.232672Z","shell.execute_reply.started":"2024-04-15T19:57:44.142399Z"}},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The first few files are of class `asbfly`:","metadata":{}},{"cell_type":"code","source":"chunk_labels[:8]","metadata":{"execution":{"iopub.execute_input":"2024-04-15T19:57:44.234546Z","iopub.status.busy":"2024-04-15T19:57:44.234353Z","iopub.status.idle":"2024-04-15T19:57:44.239514Z","shell.execute_reply":"2024-04-15T19:57:44.238788Z","shell.execute_reply.started":"2024-04-15T19:57:44.234529Z"}},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The last few files are of class `zitcis1`:","metadata":{}},{"cell_type":"code","source":"chunk_labels[-10:]","metadata":{"execution":{"iopub.execute_input":"2024-04-15T19:57:44.240381Z","iopub.status.busy":"2024-04-15T19:57:44.240210Z","iopub.status.idle":"2024-04-15T19:57:44.245570Z","shell.execute_reply":"2024-04-15T19:57:44.245020Z","shell.execute_reply.started":"2024-04-15T19:57:44.240365Z"}},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The number of `chunk_labels` should equal the `sum` of `n_chunks`:","metadata":{}},{"cell_type":"code","source":"len(chunk_labels) == sampled_df.n_chunks.sum()","metadata":{"execution":{"iopub.execute_input":"2024-04-15T19:57:44.246495Z","iopub.status.busy":"2024-04-15T19:57:44.246296Z","iopub.status.idle":"2024-04-15T19:57:44.250787Z","shell.execute_reply":"2024-04-15T19:57:44.250303Z","shell.execute_reply.started":"2024-04-15T19:57:44.246478Z"}},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Creating Chunk Tensors","metadata":{}},{"cell_type":"markdown","source":"With the filenames and class labels prepared, I can now go through and create 5-second chunks of the audio files and save them as tensors.\n\nI initially was going to create a list of all the 5-second `TensorAudio` objects with the following code, and then was going to loop through them and filenames/labels:\n\n```python\nchunks = [TensorAudio.create(path/'train_audio'/row['filename'])[...,(i * 160000):(i + 1) * 160000] \n          for _, row in sampled_df.iterrows()\n          for i in range(row['n_chunks'])]\n```\n\nHowever I ran out of memory. Instead I will use the following approach (in the code cell below) where I only load one `TensorAudio` object at a time before saving it. ChatGPT's explanation:\n\n> Comparatively, the second option is likely to have a lower maximum memory usage because it processes data iteratively, only holding one tensor in memory at a time, while the first option may consume more memory since it creates and stores all tensors at once in the list comprehension.","metadata":{}},{"cell_type":"markdown","source":"I'll create the temporary directories I need to hold tensors.","metadata":{}},{"cell_type":"code","source":"labels = sampled_df.primary_label.unique()\nlabels[:5], labels[-5:], len(labels)","metadata":{"execution":{"iopub.execute_input":"2024-04-15T19:57:44.251772Z","iopub.status.busy":"2024-04-15T19:57:44.251615Z","iopub.status.idle":"2024-04-15T19:57:44.256932Z","shell.execute_reply":"2024-04-15T19:57:44.256151Z","shell.execute_reply.started":"2024-04-15T19:57:44.251756Z"}},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_chunks_dir = '/notebooks/train_chunks'\nos.makedirs(train_chunks_dir, exist_ok=True)","metadata":{"execution":{"iopub.execute_input":"2024-04-15T21:56:50.995935Z","iopub.status.busy":"2024-04-15T21:56:50.995640Z","iopub.status.idle":"2024-04-15T21:56:51.000634Z","shell.execute_reply":"2024-04-15T21:56:50.999964Z","shell.execute_reply.started":"2024-04-15T21:56:50.995914Z"}},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for label in labels:\n    os.makedirs(train_chunks_dir + '/' + label, exist_ok=True)","metadata":{"execution":{"iopub.execute_input":"2024-04-15T21:56:52.267861Z","iopub.status.busy":"2024-04-15T21:56:52.267122Z","iopub.status.idle":"2024-04-15T21:56:52.319074Z","shell.execute_reply":"2024-04-15T21:56:52.318372Z","shell.execute_reply.started":"2024-04-15T21:56:52.267832Z"}},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"Path(train_chunks_dir).ls()","metadata":{"execution":{"iopub.execute_input":"2024-04-15T21:56:53.709648Z","iopub.status.busy":"2024-04-15T21:56:53.709102Z","iopub.status.idle":"2024-04-15T21:56:53.717257Z","shell.execute_reply":"2024-04-15T21:56:53.716556Z","shell.execute_reply.started":"2024-04-15T21:56:53.709626Z"}},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"After looking at the following code, I realized that I actually don't need `chunk_fnames` and `chunk_labels` because I can just use the iteration over `n_chunks` to construct the filenames and use `primary_label` directly to get the class label. I'll still keep the above code since it helped me visualize the process I needed to get to the code below.","metadata":{}},{"cell_type":"markdown","source":"A quick note about saving the tensors---in my first attempt I was `torch.save`-ing `TensorAudio` objects. These objects are much much larger in size on disk than normal PyTorch tensors. For example, here a 52-second `TensorAudio` object with about 1.68 million float elements, is abpit 6.71 MB on disk.","metadata":{}},{"cell_type":"code","source":"ta_full = TensorAudio.create(path/'train_audio'/'asbfly/XC785503.ogg')\nta_full.shape, ta_full.duration","metadata":{"execution":{"iopub.execute_input":"2024-04-15T21:38:46.951004Z","iopub.status.busy":"2024-04-15T21:38:46.950351Z","iopub.status.idle":"2024-04-15T21:38:46.995613Z","shell.execute_reply":"2024-04-15T21:38:46.995024Z","shell.execute_reply.started":"2024-04-15T21:38:46.950980Z"}},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"ta_full.type()","metadata":{"execution":{"iopub.execute_input":"2024-04-15T21:39:29.336631Z","iopub.status.busy":"2024-04-15T21:39:29.335943Z","iopub.status.idle":"2024-04-15T21:39:29.341398Z","shell.execute_reply":"2024-04-15T21:39:29.340899Z","shell.execute_reply.started":"2024-04-15T21:39:29.336608Z"}},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"torch.save(ta_full, \"/notebooks/ta_full.pt\")","metadata":{"execution":{"iopub.execute_input":"2024-04-15T21:37:58.242844Z","iopub.status.busy":"2024-04-15T21:37:58.242302Z","iopub.status.idle":"2024-04-15T21:37:58.256675Z","shell.execute_reply":"2024-04-15T21:37:58.256133Z","shell.execute_reply.started":"2024-04-15T21:37:58.242820Z"}},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"!wc -c /notebooks/ta_full.pt","metadata":{"execution":{"iopub.execute_input":"2024-04-15T21:38:16.823599Z","iopub.status.busy":"2024-04-15T21:38:16.822972Z","iopub.status.idle":"2024-04-15T21:38:17.290752Z","shell.execute_reply":"2024-04-15T21:38:17.290039Z","shell.execute_reply.started":"2024-04-15T21:38:16.823574Z"}},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"If I slice the first 160k samples, which is 5-seconds because the sampling rate is 32 kHz, it's disk space is still 6.71 MB!","metadata":{}},{"cell_type":"code","source":"ta_5 = ta_full[...,(0 * 160000):(0 + 1) * 160000]\nta_5.shape, ta_5.duration","metadata":{"execution":{"iopub.execute_input":"2024-04-15T21:41:14.919786Z","iopub.status.busy":"2024-04-15T21:41:14.919079Z","iopub.status.idle":"2024-04-15T21:41:14.924348Z","shell.execute_reply":"2024-04-15T21:41:14.923692Z","shell.execute_reply.started":"2024-04-15T21:41:14.919762Z"}},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"torch.save(ta_5, \"/notebooks/ta_5.pt\")","metadata":{"execution":{"iopub.execute_input":"2024-04-15T21:41:58.689458Z","iopub.status.busy":"2024-04-15T21:41:58.689194Z","iopub.status.idle":"2024-04-15T21:41:58.705416Z","shell.execute_reply":"2024-04-15T21:41:58.704626Z","shell.execute_reply.started":"2024-04-15T21:41:58.689440Z"}},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"!wc -c /notebooks/ta_5.pt","metadata":{"execution":{"iopub.execute_input":"2024-04-15T21:42:06.800410Z","iopub.status.busy":"2024-04-15T21:42:06.800066Z","iopub.status.idle":"2024-04-15T21:42:07.269878Z","shell.execute_reply":"2024-04-15T21:42:07.268945Z","shell.execute_reply.started":"2024-04-15T21:42:06.800385Z"}},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"However, if I convert it to a regular tensor, it's about 10 times smaller: 641 kB.","metadata":{}},{"cell_type":"code","source":"torch.save(ta_5.clone().detach(), \"/notebooks/ta_5_t.pt\")","metadata":{"execution":{"iopub.execute_input":"2024-04-15T21:43:11.049535Z","iopub.status.busy":"2024-04-15T21:43:11.049241Z","iopub.status.idle":"2024-04-15T21:43:11.058608Z","shell.execute_reply":"2024-04-15T21:43:11.057908Z","shell.execute_reply.started":"2024-04-15T21:43:11.049515Z"}},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"!wc -c /notebooks/ta_5_t.pt","metadata":{"execution":{"iopub.execute_input":"2024-04-15T21:43:21.249561Z","iopub.status.busy":"2024-04-15T21:43:21.248893Z","iopub.status.idle":"2024-04-15T21:43:21.714013Z","shell.execute_reply":"2024-04-15T21:43:21.713124Z","shell.execute_reply.started":"2024-04-15T21:43:21.249539Z"}},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The actual sound remains unchanged:","metadata":{}},{"cell_type":"code","source":"ta_5.show();","metadata":{"execution":{"iopub.execute_input":"2024-04-15T21:44:28.109464Z","iopub.status.busy":"2024-04-15T21:44:28.109086Z","iopub.status.idle":"2024-04-15T21:44:28.416982Z","shell.execute_reply":"2024-04-15T21:44:28.414548Z","shell.execute_reply.started":"2024-04-15T21:44:28.109443Z"}},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"TensorAudio(ta_5.clone().detach(), sr=32000).show();","metadata":{"execution":{"iopub.execute_input":"2024-04-15T21:44:57.807888Z","iopub.status.busy":"2024-04-15T21:44:57.807607Z","iopub.status.idle":"2024-04-15T21:44:58.083168Z","shell.execute_reply":"2024-04-15T21:44:58.082536Z","shell.execute_reply.started":"2024-04-15T21:44:57.807869Z"}},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"I have a total of 12826 chunks. At 641 kB per chunk, this will take up a total of 8 GB.","metadata":{}},{"cell_type":"code","source":"sampled_df.n_chunks.sum(), sampled_df.n_chunks.sum() * 641000 / 1e9","metadata":{"execution":{"iopub.execute_input":"2024-04-15T21:47:13.186173Z","iopub.status.busy":"2024-04-15T21:47:13.185608Z","iopub.status.idle":"2024-04-15T21:47:13.190419Z","shell.execute_reply":"2024-04-15T21:47:13.189912Z","shell.execute_reply.started":"2024-04-15T21:47:13.186151Z"}},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"That's still a lot of data just for a quick-and-dirty setup. I'll limit `n_chunks` to 5 for each audio file in `sample_df`. ","metadata":{}},{"cell_type":"code","source":"sampled_df['n_chunks'] = sampled_df['n_chunks'].apply(lambda x: x if x<5 else 5)","metadata":{"execution":{"iopub.execute_input":"2024-04-15T21:50:07.689310Z","iopub.status.busy":"2024-04-15T21:50:07.688641Z","iopub.status.idle":"2024-04-15T21:50:07.694034Z","shell.execute_reply":"2024-04-15T21:50:07.693375Z","shell.execute_reply.started":"2024-04-15T21:50:07.689286Z"}},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sampled_df['n_chunks'].value_counts()","metadata":{"execution":{"iopub.execute_input":"2024-04-15T21:50:23.420739Z","iopub.status.busy":"2024-04-15T21:50:23.420466Z","iopub.status.idle":"2024-04-15T21:50:23.428564Z","shell.execute_reply":"2024-04-15T21:50:23.428066Z","shell.execute_reply.started":"2024-04-15T21:50:23.420720Z"}},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"For now, I'll see if I can remove files that have zero five-second chunks, as long as I still have all 182 of my bird classes.","metadata":{}},{"cell_type":"code","source":"len(sampled_df.query(\"n_chunks > 0\").primary_label.unique())","metadata":{"execution":{"iopub.execute_input":"2024-04-15T21:52:40.283282Z","iopub.status.busy":"2024-04-15T21:52:40.282998Z","iopub.status.idle":"2024-04-15T21:52:40.292786Z","shell.execute_reply":"2024-04-15T21:52:40.292190Z","shell.execute_reply.started":"2024-04-15T21:52:40.283262Z"}},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Great! I can remove all files less than 5 seconds long without losing any of my labels.","metadata":{}},{"cell_type":"code","source":"sampled_df_final = sampled_df.query(\"n_chunks > 0\")\nsampled_df_final.n_chunks.sum()","metadata":{"execution":{"iopub.execute_input":"2024-04-15T21:53:43.166203Z","iopub.status.busy":"2024-04-15T21:53:43.165693Z","iopub.status.idle":"2024-04-15T21:53:43.174248Z","shell.execute_reply":"2024-04-15T21:53:43.173547Z","shell.execute_reply.started":"2024-04-15T21:53:43.166180Z"}},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"6158 chunks, at 641 kB per chunk, is a total of 3.94 GB.","metadata":{}},{"cell_type":"code","source":"6158*641000/1e9","metadata":{"execution":{"iopub.execute_input":"2024-04-15T21:54:23.211470Z","iopub.status.busy":"2024-04-15T21:54:23.210802Z","iopub.status.idle":"2024-04-15T21:54:23.216171Z","shell.execute_reply":"2024-04-15T21:54:23.215307Z","shell.execute_reply.started":"2024-04-15T21:54:23.211444Z"}},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"len(sampled_df_final)","metadata":{"execution":{"iopub.execute_input":"2024-04-15T21:57:29.897763Z","iopub.status.busy":"2024-04-15T21:57:29.897420Z","iopub.status.idle":"2024-04-15T21:57:29.903192Z","shell.execute_reply":"2024-04-15T21:57:29.902422Z","shell.execute_reply.started":"2024-04-15T21:57:29.897740Z"}},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for _, row in sampled_df_final.iterrows():\n    for i in range(row['n_chunks']):\n        # create chunk tensor file name stem\n        end_time = (i + 1) * 5\n        fn_stem = f\"{row['filename'].split('/')[1].split('.')[0]}_{end_time}\" \n        \n        # create full file path for chunk tensor\n        fn = Path(train_chunks_dir)/row['primary_label']/(fn_stem + '.pt')\n        \n        # load a chunk of the TensorAudio\n        ta = TensorAudio.create(path/'train_audio'/row['filename'])[...,(i * 160000):(i + 1) * 160000]\n        \n        # convert it to regular tenros\n        ta = ta.clone().detach()\n        \n        # save it\n        torch.save(ta, fn)","metadata":{"execution":{"iopub.execute_input":"2024-04-15T22:00:33.883932Z","iopub.status.busy":"2024-04-15T22:00:33.883642Z","iopub.status.idle":"2024-04-15T22:04:08.747323Z","shell.execute_reply":"2024-04-15T22:04:08.746546Z","shell.execute_reply.started":"2024-04-15T22:00:33.883911Z"}},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"I'll check the total size of the tensors:","metadata":{}},{"cell_type":"code","source":"!du -hs /notebooks/train_chunks","metadata":{"execution":{"iopub.execute_input":"2024-04-15T22:07:06.054177Z","iopub.status.busy":"2024-04-15T22:07:06.053343Z","iopub.status.idle":"2024-04-15T22:07:06.643486Z","shell.execute_reply":"2024-04-15T22:07:06.642672Z","shell.execute_reply.started":"2024-04-15T22:07:06.054150Z"}},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"3.7 GB, close to what I expected. Whew!","metadata":{}},{"cell_type":"code","source":"import time\ntime.sleep(6000)","metadata":{"execution":{"iopub.status.busy":"2024-04-15T20:10:48.094776Z","iopub.status.idle":"2024-04-15T20:10:48.095003Z","shell.execute_reply":"2024-04-15T20:10:48.094912Z","shell.execute_reply.started":"2024-04-15T20:10:48.094902Z"}},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Creating `DataLoaders`","metadata":{}},{"cell_type":"markdown","source":"The `DataBlock` will have to achieve the following:\n\n- Get the filepath to each tensor\n- Convert it to a `TensorAudio` object\n- Pass it through `PowerToDB`\n- Use the parent folder name as the label","metadata":{}},{"cell_type":"markdown","source":"I'll use the `PowerToDB` transform I created in my [previous notebook](https://www.kaggle.com/code/vishalbakshi/birdclef2024-getting-started-with-fastxtend-audio/edit) with a couple of changes:\n\nInstead of using `torchaudio.transforms.MelSpectrogram` I'll use fastxtend's `MelSpectrogram`, and I'll set the default sample rate to 32000 to match the training data.\n\nI also had to add `order = 75` at the top since of the class definition. Initially I was passing the `DataBlock` full files and then using `RandomCropPad` to crop them to 5 seconds. When I checked the pipeline with `DataBlock.summary(df)` I saw that even though I had placed `PowerToDB` after `RandomCropPad` in the `Pipeline`, the items were passing through `PowerToDB` first and THEN were being passed through `RandomCropPad` and so the number of samples were not consistent across the items (I was getting an error when running `DataBlock.dataloaders(df)` that the tensors' shapes didn't match when trying to create a match). After adding `order = 75` the order of the `Pipeline` was correct.","metadata":{}},{"cell_type":"code","source":"class PowerToDB(Transform):\n    order = 75\n    def __init__(self, sr=32000, n_fft=512): \n        self.sr = sr\n        self.n_fft=n_fft\n        self.mel = MelSpectrogram(sample_rate=self.sr, n_fft=self.n_fft)\n    def encodes(self, x:TensorAudio):\n        return TensorMelSpec.create(librosa.power_to_db(self.mel(x).cpu()), settings={})","metadata":{"execution":{"iopub.execute_input":"2024-04-15T22:11:16.032279Z","iopub.status.busy":"2024-04-15T22:11:16.031674Z","iopub.status.idle":"2024-04-15T22:11:16.036959Z","shell.execute_reply":"2024-04-15T22:11:16.036379Z","shell.execute_reply.started":"2024-04-15T22:11:16.032256Z"}},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"I'll use a splitter that randomly assigns 20% of the rows to the validation set. In future iterations, I'll be more thoughtful about the validation split as in this case, potentially one species could dominate the validation set instead of a proportional mix.","metadata":{}},{"cell_type":"markdown","source":"In the first `DataLoaders` object I create, I'll use the filename as the label so I can confirm the spectrogram is being made correctly. I'll also need to define a `LoadTensorAudio` transform that takes a filepath, loads it as a tensor and creates a `TensorAudio` object. I can't just use `TensorAudio.create` since it is expecting an audio file as input, whereas I have tensors.","metadata":{}},{"cell_type":"code","source":"class LoadTensorAudio(Transform):\n    def __init__(self, sr=32000):\n        self.sr = sr\n    def encodes(self, x:Path):\n        x = torch.load(x)\n        return TensorAudio(x, sr=self.sr)","metadata":{"execution":{"iopub.execute_input":"2024-04-15T22:23:06.636138Z","iopub.status.busy":"2024-04-15T22:23:06.635227Z","iopub.status.idle":"2024-04-15T22:23:06.640035Z","shell.execute_reply":"2024-04-15T22:23:06.639396Z","shell.execute_reply.started":"2024-04-15T22:23:06.636114Z"}},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"auds = DataBlock(blocks = (TransformBlock, CategoryBlock),  \n                 get_items=get_files,\n                 item_tfms = [LoadTensorAudio, PowerToDB(sr=32000)], \n                 splitter=RandomSplitter(valid_pct=0.2, seed=42),\n                 get_y = lambda o: o.parent.name + '/' + o.stem)","metadata":{"execution":{"iopub.execute_input":"2024-04-15T22:25:10.366378Z","iopub.status.busy":"2024-04-15T22:25:10.365787Z","iopub.status.idle":"2024-04-15T22:25:10.371089Z","shell.execute_reply":"2024-04-15T22:25:10.370653Z","shell.execute_reply.started":"2024-04-15T22:25:10.366357Z"}},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"dls = auds.dataloaders(Path(train_chunks_dir), bs=64)\ndls.show_batch(figsize=(10, 5))","metadata":{"execution":{"iopub.execute_input":"2024-04-15T22:25:10.986422Z","iopub.status.busy":"2024-04-15T22:25:10.985913Z","iopub.status.idle":"2024-04-15T22:25:14.300356Z","shell.execute_reply":"2024-04-15T22:25:14.299713Z","shell.execute_reply.started":"2024-04-15T22:25:10.986397Z"}},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"So far, it looks good! I'll see if I can recreate the first spectrogram in the `show_batch` output to check that my transform is working correctly.","metadata":{}},{"cell_type":"code","source":"# create a TensorAudio object\nta_check = torch.load(Path(train_chunks_dir)/'houspa'/'XC479870_25.pt')\nta_check = TensorAudio(ta_check, sr=32000)\nta_check.shape, ta_check.duration","metadata":{"execution":{"iopub.execute_input":"2024-04-15T22:27:08.883984Z","iopub.status.busy":"2024-04-15T22:27:08.883359Z","iopub.status.idle":"2024-04-15T22:27:08.889448Z","shell.execute_reply":"2024-04-15T22:27:08.889004Z","shell.execute_reply.started":"2024-04-15T22:27:08.883958Z"}},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# create scaled MelSpectrogram\nscaled_mel = librosa.power_to_db(MelSpectrogram(sample_rate=32000, n_fft=512)(ta_check))","metadata":{"execution":{"iopub.execute_input":"2024-04-15T22:27:10.160692Z","iopub.status.busy":"2024-04-15T22:27:10.159908Z","iopub.status.idle":"2024-04-15T22:27:10.293514Z","shell.execute_reply":"2024-04-15T22:27:10.293022Z","shell.execute_reply.started":"2024-04-15T22:27:10.160668Z"}},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"TensorMelSpec.create(scaled_mel, settings={}).shape","metadata":{"execution":{"iopub.execute_input":"2024-04-15T22:27:10.735319Z","iopub.status.busy":"2024-04-15T22:27:10.734514Z","iopub.status.idle":"2024-04-15T22:27:10.740323Z","shell.execute_reply":"2024-04-15T22:27:10.739702Z","shell.execute_reply.started":"2024-04-15T22:27:10.735287Z"}},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"TensorMelSpec.create(scaled_mel, settings={}).show();","metadata":{"execution":{"iopub.execute_input":"2024-04-15T22:27:11.128071Z","iopub.status.busy":"2024-04-15T22:27:11.127282Z","iopub.status.idle":"2024-04-15T22:27:11.310672Z","shell.execute_reply":"2024-04-15T22:27:11.310230Z","shell.execute_reply.started":"2024-04-15T22:27:11.128049Z"}},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Looks good! The aspect ratio is different but the x- and y-axis intervals and overall spectrogram signal looks the same as the `show_batch` output. I'll now create a `DataLoaders` with the correct class label.","metadata":{}},{"cell_type":"code","source":"auds = DataBlock(blocks = (TransformBlock, CategoryBlock),  \n                 get_items=get_files,\n                 item_tfms = [LoadTensorAudio, PowerToDB(sr=32000)], \n                 splitter=RandomSplitter(valid_pct=0.2, seed=42),\n                 get_y = lambda o: o.parent.name)","metadata":{"execution":{"iopub.execute_input":"2024-04-15T22:27:36.964802Z","iopub.status.busy":"2024-04-15T22:27:36.964026Z","iopub.status.idle":"2024-04-15T22:27:37.013291Z","shell.execute_reply":"2024-04-15T22:27:37.012761Z","shell.execute_reply.started":"2024-04-15T22:27:36.964778Z"}},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"dls = auds.dataloaders(Path(train_chunks_dir), bs=64)\ndls.show_batch(figsize=(10, 5))","metadata":{"execution":{"iopub.execute_input":"2024-04-15T22:28:00.905090Z","iopub.status.busy":"2024-04-15T22:28:00.904432Z","iopub.status.idle":"2024-04-15T22:28:05.993825Z","shell.execute_reply":"2024-04-15T22:28:05.993246Z","shell.execute_reply.started":"2024-04-15T22:28:00.905067Z"}},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Looks good! I'll train it for 10 epochs.","metadata":{}},{"cell_type":"code","source":"learn = vision_learner(\n    dls, \n    resnet18,\n    n_in=1,\n    loss_func=CrossEntropyLossFlat(),\n    metrics=[accuracy])\n\nlearn.fine_tune(10)","metadata":{"execution":{"iopub.execute_input":"2024-04-15T22:28:52.162629Z","iopub.status.busy":"2024-04-15T22:28:52.162311Z","iopub.status.idle":"2024-04-15T22:31:15.678515Z","shell.execute_reply":"2024-04-15T22:31:15.677778Z","shell.execute_reply.started":"2024-04-15T22:28:52.162607Z"}},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Hooray! It trains!! And it trains in under 3 minutes. The accuracy is not great, but not terrible at 67%. I'll export it and then create a corresponding [Submit] notebook for inference.\n\nI hope you enjoyed this notebook! If you did, please upvote (thanks).","metadata":{}},{"cell_type":"code","source":"learn.model_dir = '/notebooks'","metadata":{"execution":{"iopub.execute_input":"2024-04-15T22:31:42.776777Z","iopub.status.busy":"2024-04-15T22:31:42.776133Z","iopub.status.idle":"2024-04-15T22:31:42.779966Z","shell.execute_reply":"2024-04-15T22:31:42.779352Z","shell.execute_reply.started":"2024-04-15T22:31:42.776754Z"}},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"learn.save('birdclef24_resnet18_v2', with_opt=False)","metadata":{"execution":{"iopub.execute_input":"2024-04-15T22:31:46.665807Z","iopub.status.busy":"2024-04-15T22:31:46.665122Z","iopub.status.idle":"2024-04-15T22:31:46.753381Z","shell.execute_reply":"2024-04-15T22:31:46.752700Z","shell.execute_reply.started":"2024-04-15T22:31:46.665781Z"}},"execution_count":null,"outputs":[]}]}