{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Notebook Overview\nThe purpose of this notebook is to establish a data reformatting pipeline for eventual model training using Keras and the TensorFlow API. Prior to this notebook, I created a preprocessing pipeline that would convert the input '.ogg' files into spectrogram representations. This notebook, however, will put emphasis on refactoring dataframe objects to consolidate input-output data into single rows.","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19"}},{"cell_type":"markdown","source":"### General Format\nA given row will contain the following:\n* A numerical representation of the input data (via spectrogram)\n* Columns of all species data and their relative 'exsistance chance' in the audio file\n\nA given row might look like this:\n\"[input audio]-[species 1: chance],[species 2: chance],...[species n: chance]\"","metadata":{}},{"cell_type":"code","source":"#importing dependencies\nimport os\nimport pandas as pd\nimport numpy as np","metadata":{"execution":{"iopub.status.busy":"2023-03-11T02:18:39.765080Z","iopub.execute_input":"2023-03-11T02:18:39.765566Z","iopub.status.idle":"2023-03-11T02:18:39.772690Z","shell.execute_reply.started":"2023-03-11T02:18:39.765527Z","shell.execute_reply":"2023-03-11T02:18:39.771088Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"metadata = pd.read_csv('/kaggle/input/birdclef-2023/train_metadata.csv')\nmetadata.head(1)","metadata":{"execution":{"iopub.status.busy":"2023-03-11T02:18:39.782738Z","iopub.execute_input":"2023-03-11T02:18:39.784387Z","iopub.status.idle":"2023-03-11T02:18:39.872804Z","shell.execute_reply.started":"2023-03-11T02:18:39.784324Z","shell.execute_reply":"2023-03-11T02:18:39.871375Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#list to contain all found species in metadata dataframe object\nspecies = metadata['primary_label'].unique().tolist()","metadata":{"execution":{"iopub.status.busy":"2023-03-11T02:18:39.875475Z","iopub.execute_input":"2023-03-11T02:18:39.876872Z","iopub.status.idle":"2023-03-11T02:18:39.884673Z","shell.execute_reply.started":"2023-03-11T02:18:39.876804Z","shell.execute_reply":"2023-03-11T02:18:39.883601Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#adding full filepath column to metadata dataframe\nmetadata['filename'] = metadata['filename'].apply(lambda x : \"/kaggle/input/birdclef-2023/train_audio/\" + x)","metadata":{"execution":{"iopub.status.busy":"2023-03-11T02:18:39.885900Z","iopub.execute_input":"2023-03-11T02:18:39.886561Z","iopub.status.idle":"2023-03-11T02:18:39.904258Z","shell.execute_reply.started":"2023-03-11T02:18:39.886523Z","shell.execute_reply":"2023-03-11T02:18:39.902699Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"metadata.columns","metadata":{"execution":{"iopub.status.busy":"2023-03-11T02:18:39.906837Z","iopub.execute_input":"2023-03-11T02:18:39.907383Z","iopub.status.idle":"2023-03-11T02:18:39.920401Z","shell.execute_reply.started":"2023-03-11T02:18:39.907337Z","shell.execute_reply":"2023-03-11T02:18:39.918543Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#dropping other columns from metadata\nmetadata = metadata.drop(['secondary_labels', 'type', 'latitude', 'longitude',\n       'scientific_name', 'common_name', 'author', 'license', 'rating', 'url'], axis=1)","metadata":{"execution":{"iopub.status.busy":"2023-03-11T02:18:39.922512Z","iopub.execute_input":"2023-03-11T02:18:39.923046Z","iopub.status.idle":"2023-03-11T02:18:39.933348Z","shell.execute_reply.started":"2023-03-11T02:18:39.922987Z","shell.execute_reply":"2023-03-11T02:18:39.932075Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Preprocessing Checkpoint\nAt the current time, there is a species associated with a given filename. Next, we will store the numerical data associated with the audio file in that filepath in a column named 'spectrogram'. This will eventually replace the filepath column once preprocessing has been complete.","metadata":{"execution":{"iopub.status.busy":"2023-03-10T17:16:23.858661Z","iopub.execute_input":"2023-03-10T17:16:23.859204Z","iopub.status.idle":"2023-03-10T17:16:23.873935Z","shell.execute_reply.started":"2023-03-10T17:16:23.859141Z","shell.execute_reply":"2023-03-10T17:16:23.872334Z"}}},{"cell_type":"code","source":"#importing dependancies for '.ogg' -> spectrograme given filepath\nimport librosa\nfrom scipy import signal\nfrom matplotlib import pyplot as plt","metadata":{"execution":{"iopub.status.busy":"2023-03-11T02:18:39.935176Z","iopub.execute_input":"2023-03-11T02:18:39.936357Z","iopub.status.idle":"2023-03-11T02:18:39.947510Z","shell.execute_reply.started":"2023-03-11T02:18:39.936303Z","shell.execute_reply":"2023-03-11T02:18:39.946214Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# will convert an .ogg file to a waveform\ndef ogg_to_wave(filename):\n    ogg, sample_rate = librosa.load(filename)\n    int16 = (ogg * 32767).astype(np.int16)\n    return int16, sample_rate","metadata":{"execution":{"iopub.status.busy":"2023-03-11T02:18:39.949108Z","iopub.execute_input":"2023-03-11T02:18:39.949935Z","iopub.status.idle":"2023-03-11T02:18:39.958594Z","shell.execute_reply.started":"2023-03-11T02:18:39.949895Z","shell.execute_reply":"2023-03-11T02:18:39.957399Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#will take in created waveform and create a spectogram matrix for neural network input layer\ndef wave_to_spec(waveform, sample_rate):\n    freq, time, spectrogram = signal.spectrogram(waveform, sample_rate, window='hann', nperseg=256, noverlap=128)\n    return freq, time, spectrogram","metadata":{"execution":{"iopub.status.busy":"2023-03-11T02:18:39.959925Z","iopub.execute_input":"2023-03-11T02:18:39.960374Z","iopub.status.idle":"2023-03-11T02:18:39.972729Z","shell.execute_reply.started":"2023-03-11T02:18:39.960330Z","shell.execute_reply":"2023-03-11T02:18:39.971618Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#function assembly for preprocessing data\ndef audio_preprocessing(filepath):\n    waveform, sample_rate = ogg_to_wave(filepath)\n    freq, time, spectrogram = wave_to_spec(waveform, sample_rate)\n    return freq, time, spectrogram","metadata":{"execution":{"iopub.status.busy":"2023-03-11T02:18:39.974197Z","iopub.execute_input":"2023-03-11T02:18:39.975217Z","iopub.status.idle":"2023-03-11T02:18:39.987728Z","shell.execute_reply.started":"2023-03-11T02:18:39.975125Z","shell.execute_reply":"2023-03-11T02:18:39.986625Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# test ogg_to_wave function\nwave_test, wave_test_sample_rate = ogg_to_wave(\"/kaggle/input/birdclef-2023/train_audio/afrgrp1/XC126598.ogg\")\nplt.plot(wave_test)\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-03-11T02:18:39.991087Z","iopub.execute_input":"2023-03-11T02:18:39.991626Z","iopub.status.idle":"2023-03-11T02:18:40.571668Z","shell.execute_reply.started":"2023-03-11T02:18:39.991583Z","shell.execute_reply":"2023-03-11T02:18:40.570585Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#test wave_to_spec function\nf, t, s = wave_to_spec(wave_test, wave_test_sample_rate)\n\n#plotting out spectrogram\nplt.pcolormesh(t, f, 10*np.log10(s), cmap='viridis')\nplt.xlabel(\"Time(s)\")\nplt.ylabel(\"Frequency (Hz)\")\nplt.title(\"Audio Spectrogram Representation\")\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-03-11T02:18:40.573307Z","iopub.execute_input":"2023-03-11T02:18:40.573962Z","iopub.status.idle":"2023-03-11T02:18:41.473448Z","shell.execute_reply.started":"2023-03-11T02:18:40.573922Z","shell.execute_reply":"2023-03-11T02:18:41.472044Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#test of full audio preprocessing function pipeline\nf, t, s = audio_preprocessing(\"/kaggle/input/birdclef-2023/train_audio/afpwag1/XC131306.ogg\")\n\nplt.pcolormesh(t, f, 10*np.log10(s), cmap='viridis')\nplt.xlabel(\"Time(s)\")\nplt.ylabel(\"Frequency (Hz)\")\nplt.title(\"Audio Spectrogram Representation\")\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-03-11T02:18:41.475059Z","iopub.execute_input":"2023-03-11T02:18:41.476392Z","iopub.status.idle":"2023-03-11T02:18:42.298781Z","shell.execute_reply.started":"2023-03-11T02:18:41.476335Z","shell.execute_reply":"2023-03-11T02:18:42.297187Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Preprocessing Checkpoint\nNow at this point, there is a defined datapipeline that will take a filepath and return its spectrogram representation. The next step is to create a dataframe such that the original formatting described in the project overview will be satisfied.\n\nRefresh of desired input format:\n\n\"[input audio]-[species 1: chance],[species 2: chance],...[species n: chance]\"","metadata":{}},{"cell_type":"code","source":"#testing pipeline on one audio file from each species (will allow one-hot encoding)\ntest = metadata.drop_duplicates(subset=[\"primary_label\"], keep=False)\ntest.head()","metadata":{"execution":{"iopub.status.busy":"2023-03-11T02:18:42.300382Z","iopub.execute_input":"2023-03-11T02:18:42.300758Z","iopub.status.idle":"2023-03-11T02:18:42.318452Z","shell.execute_reply.started":"2023-03-11T02:18:42.300721Z","shell.execute_reply":"2023-03-11T02:18:42.316831Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#replacing filename with spectrogram representation\ntest['spectrogram'] = test['filename'].apply(lambda x : audio_preprocessing(x)[2])\ntest.drop('filename', axis=1, inplace=True)","metadata":{"execution":{"iopub.status.busy":"2023-03-11T02:18:42.320380Z","iopub.execute_input":"2023-03-11T02:18:42.320758Z","iopub.status.idle":"2023-03-11T02:18:42.854392Z","shell.execute_reply.started":"2023-03-11T02:18:42.320722Z","shell.execute_reply":"2023-03-11T02:18:42.852955Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test.head()","metadata":{"execution":{"iopub.status.busy":"2023-03-11T02:18:42.856064Z","iopub.execute_input":"2023-03-11T02:18:42.856587Z","iopub.status.idle":"2023-03-11T02:18:43.421204Z","shell.execute_reply.started":"2023-03-11T02:18:42.856522Z","shell.execute_reply":"2023-03-11T02:18:43.419646Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test.reset_index(inplace=True)\n#one-hot encoding species column\nspecies = pd.get_dummies(test['primary_label']).astype('float64')\n#dropping primary_label column (will be replaced with OH encoded columns)\ntest = test.drop(['primary_label'], axis=1)\ntest = test.join(species)\ntest.head()","metadata":{"execution":{"iopub.status.busy":"2023-03-11T02:18:43.422909Z","iopub.execute_input":"2023-03-11T02:18:43.424058Z","iopub.status.idle":"2023-03-11T02:18:43.978214Z","shell.execute_reply.started":"2023-03-11T02:18:43.424012Z","shell.execute_reply":"2023-03-11T02:18:43.977245Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Preprocessing Checkpoint\nNow that we have a way to create the desired input dataframe, let's establish a final function to tie it all together!\n\n\n#### Applied to each row\n(Look at row from 'train_metadata.csv') -> (get audiofile path) -> (convert '.ogg' to waveform) -> (convert waveform to spectrogram) -> (add spectrogram to given row)\n\n#### Applied once all rows have been audio-encoded\n-> (delete other rows) -> (encode species data)","metadata":{}},{"cell_type":"code","source":"#will take in .csv filepath (formatting expected to reflect that described earlier)\ndef preprocessing(dataframe):\n    \n    #creating dataframe\n    df = pd.read_csv(dataframe).head(5)\n    \n    #creating spectrogram column\n    file_prefix = \"/kaggle/input/birdclef-2023/train_audio/\"\n    df['input'] = df['filename'].apply(lambda x : audio_preprocessing(file_prefix + x)[2])\n    \n    #dropping other columns\n    df = df[['input', 'primary_label']]\n    \n    #one-hot encoding primary_label column\n    species = pd.get_dummies(df['primary_label']).astype('float64')\n    df = df.drop('primary_label', axis=1)\n    df = df.join(species)\n    \n    return df","metadata":{"execution":{"iopub.status.busy":"2023-03-11T02:18:43.979588Z","iopub.execute_input":"2023-03-11T02:18:43.979964Z","iopub.status.idle":"2023-03-11T02:18:43.987999Z","shell.execute_reply.started":"2023-03-11T02:18:43.979930Z","shell.execute_reply":"2023-03-11T02:18:43.986220Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sep_test_1 = preprocessing(\"/kaggle/input/birdclef-2023/train_metadata.csv\")\nsep_test_1.head()","metadata":{"execution":{"iopub.status.busy":"2023-03-11T02:18:43.989795Z","iopub.execute_input":"2023-03-11T02:18:43.990941Z","iopub.status.idle":"2023-03-11T02:18:45.085899Z","shell.execute_reply.started":"2023-03-11T02:18:43.990877Z","shell.execute_reply":"2023-03-11T02:18:45.084413Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}