{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Overview\n### Objective\nThe purpose of this notebook is to develop a neural network capable of predicting the existance of certain bird calls in an audio file.\n\n### Purpose\nThe purpose of this model is to provide a more efficient alternative to traditional methods of population sampling, which is both time-intensive and costly. This approach would allow researchers to place microphones in areas of a forest where biodiversity might be threatened, and process that data through an algorithm to determine which birds are in what areas.\n\n### Model Evaluation\nThe model will output a vector of likelyhoods that a given species is present in the audio file, where each input in the vector [c1, c2, ..., cn] will represent a given species' likelyhood of occurance. To grade this output, the cmAP metric will be used. This is a derivative of the macro-averaged average precision score as implemented by scikit-learn.\n\n### Data Organization\nThe relevant training data is split into two types: metadata and audio file. The metadata is stored as a .csv file where each row represents the metadata for a .ogg audio file. this .csv file contains things like geolocation of the recording, the type of call/song recorded, and most importantly the species of bird recorded.\n\n### Plan of Attack\nThe general outline of development will be as follows:\n1. Exploratory Data Analysis\n2. Audio Data Preprocessing\n3. Data Transforming\n4. CNN Model Creation\n5. Hyperparameter Tuning\n6. Evaluation","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19"}},{"cell_type":"markdown","source":"# 1. Exploratory Data Analysis","metadata":{}},{"cell_type":"markdown","source":"### 1.1 'train_metadata.csv' EDA","metadata":{}},{"cell_type":"code","source":"# Importing EDA dependencies\nimport pandas as pd\nimport numpy as np\nfrom pandas_profiling import ProfileReport\nimport matplotlib.pyplot as plt\nimport seaborn as sns","metadata":{"execution":{"iopub.status.busy":"2023-03-29T00:59:38.707874Z","iopub.execute_input":"2023-03-29T00:59:38.708385Z","iopub.status.idle":"2023-03-29T00:59:41.886535Z","shell.execute_reply.started":"2023-03-29T00:59:38.708339Z","shell.execute_reply":"2023-03-29T00:59:41.884486Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Reading in metadata file\nmetadata = pd.read_csv('/kaggle/input/birdclef-2023/train_metadata.csv')","metadata":{"execution":{"iopub.status.busy":"2023-03-29T00:59:41.891103Z","iopub.execute_input":"2023-03-29T00:59:41.891806Z","iopub.status.idle":"2023-03-29T00:59:42.030758Z","shell.execute_reply.started":"2023-03-29T00:59:41.891743Z","shell.execute_reply":"2023-03-29T00:59:42.028647Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Generating EDA report\nreport = ProfileReport(metadata, title='BirdCLEF 2023 Training Metadata')\nreport","metadata":{"execution":{"iopub.status.busy":"2023-03-29T00:59:42.032972Z","iopub.execute_input":"2023-03-29T00:59:42.033423Z","iopub.status.idle":"2023-03-29T00:59:51.228688Z","shell.execute_reply.started":"2023-03-29T00:59:42.033384Z","shell.execute_reply":"2023-03-29T00:59:51.226425Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### EDA Findings\n1. The training dataset has 264 unique species\n2. All rows connect to a unique audio file\n3. There are ~17,000 audio files in the dataset\n4. 'primary_label' = Species Name\n5. The most common species have 500 audio files for training (3.0% of dataset)\n6. The only useful columns will be 'primary_label' and 'filename'\n7. The '.ogg' files were standardized to a 32kHz sample rate","metadata":{}},{"cell_type":"markdown","source":"# 2. Audio Preprocessing Function\nDue to the large amount of data needing to be processed, I'm going to impliment a data pipeline using the Keras API (keras.data.Dataset, specifically). This will allow batch processing to natively occur and will negate the possible RAM/HDD bottlenecks from trying to convert all files to spectrogams using .apply() method on a dataframe, for example.","metadata":{}},{"cell_type":"markdown","source":"### 2.1 Preparing Dataframe\nThe dataframe will need to be transformed such that each row is converted into a tuple, where the first index is the filename (to be converted to a spectrogram) and the second index is all of the species labels and their probability of occurance.","metadata":{}},{"cell_type":"code","source":"# Full filepath (input)\nX = np.asarray(metadata['filename'].apply(lambda x : f'/kaggle/input/birdclef-2023/train_audio/{x}'))\n\n# One-hot encoded labels (output)\ny = np.asarray(pd.get_dummies(metadata['primary_label']))","metadata":{"execution":{"iopub.status.busy":"2023-03-29T00:59:51.231573Z","iopub.execute_input":"2023-03-29T00:59:51.232034Z","iopub.status.idle":"2023-03-29T00:59:51.255363Z","shell.execute_reply.started":"2023-03-29T00:59:51.231985Z","shell.execute_reply":"2023-03-29T00:59:51.252763Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### 2.2 Dataloading Function\nThis will take in a tuple (input, output) and return the same (intput, output) tuple, except the input will have been coverted from a filepath to a tensorflow waveform object.","metadata":{}},{"cell_type":"code","source":"# Dataloading dependencies\nimport librosa\nimport tensorflow as tf\nimport io","metadata":{"execution":{"iopub.status.busy":"2023-03-29T00:59:51.256994Z","iopub.execute_input":"2023-03-29T00:59:51.257400Z","iopub.status.idle":"2023-03-29T01:00:03.331792Z","shell.execute_reply.started":"2023-03-29T00:59:51.257352Z","shell.execute_reply":"2023-03-29T01:00:03.328926Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def load_audio_file(file_path):\n    # Read audio file contents into a tensor\n    audio_binary = tf.io.read_file(file_path)\n    \n    # Decode audio binary into a numpy array\n    audio, sr = librosa.load(io.BytesIO(audio_binary.numpy()), sr=16000, mono=True)\n    \n    audio = audio * 32768.0\n    audio = tf.cast(audio, tf.int16)\n\n    return audio","metadata":{"execution":{"iopub.status.busy":"2023-03-29T01:00:03.334893Z","iopub.execute_input":"2023-03-29T01:00:03.336841Z","iopub.status.idle":"2023-03-29T01:00:03.347597Z","shell.execute_reply.started":"2023-03-29T01:00:03.336783Z","shell.execute_reply":"2023-03-29T01:00:03.345124Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Testing dataloading function\nX_test= load_audio_file(X[0])","metadata":{"execution":{"iopub.status.busy":"2023-03-29T01:00:03.350052Z","iopub.execute_input":"2023-03-29T01:00:03.350704Z","iopub.status.idle":"2023-03-29T01:00:06.612406Z","shell.execute_reply.started":"2023-03-29T01:00:03.350652Z","shell.execute_reply":"2023-03-29T01:00:06.609331Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Printing datatype and shape\nprint(type(X_test))\nprint(X_test.shape)","metadata":{"execution":{"iopub.status.busy":"2023-03-29T01:00:06.615403Z","iopub.execute_input":"2023-03-29T01:00:06.616715Z","iopub.status.idle":"2023-03-29T01:00:06.627514Z","shell.execute_reply.started":"2023-03-29T01:00:06.616657Z","shell.execute_reply":"2023-03-29T01:00:06.625015Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### 2.3 Plotting Waveform","metadata":{}},{"cell_type":"code","source":"plt.plot(X_test)\nplt.title('Bird Call Waveform (16 kHz Downsample)')\nplt.xlabel('Time (16,000 = 1 second)')\nplt.ylabel('Waveform Amplitude')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-03-29T01:00:06.629332Z","iopub.execute_input":"2023-03-29T01:00:06.629813Z","iopub.status.idle":"2023-03-29T01:00:06.989824Z","shell.execute_reply.started":"2023-03-29T01:00:06.629759Z","shell.execute_reply":"2023-03-29T01:00:06.988519Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### 2.4 Preprocessing Function\nNext, I'll define a function that takes the now transformed (input, output) tuple and again transform the input index to be a spectrogram as compared to a waveform. This will effectively turn the audio file into an image that can be fed into the CNN.\n\nThis will be the function called by the .map() method from the Keras Dataset object to convert a given row of data into what we need it to be before training a model on it.","metadata":{}},{"cell_type":"code","source":"def preprocess(filepath, labels):\n    # Loading in audio\n    waveform = tf.py_function(load_audio_file, [filepath], tf.int16)\n    \n    # Normalize waveform to [-1, 1]\n    waveform = tf.cast(waveform, tf.float32) / 32768.0\n    \n    # Getting first 3 seconds of clip\n    waveform = waveform[:48000]\n    \n    # Add zero padding\n    zero_padding = tf.zeros([48000] - tf.shape(waveform), dtype=tf.float32)\n    wav = tf.concat([zero_padding, waveform], 0)\n    \n    # Spectrogram\n    stft = tf.signal.stft(wav, frame_length=320, frame_step=32)\n    spectrogram = tf.abs(stft)\n    spectrogram = tf.expand_dims(spectrogram, axis=2)\n    \n    return spectrogram, labels","metadata":{"execution":{"iopub.status.busy":"2023-03-29T01:00:06.995302Z","iopub.execute_input":"2023-03-29T01:00:06.995902Z","iopub.status.idle":"2023-03-29T01:00:07.008467Z","shell.execute_reply.started":"2023-03-29T01:00:06.995836Z","shell.execute_reply":"2023-03-29T01:00:07.006120Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### 2.5 Testing Preprocessing Function\nUsing one example row from the X and y series, I'm going to verify that the preprocess function completes all needed transformations of input data before creating the model.","metadata":{"execution":{"iopub.status.busy":"2023-03-27T19:46:05.030032Z","iopub.execute_input":"2023-03-27T19:46:05.030505Z","iopub.status.idle":"2023-03-27T19:46:05.390835Z","shell.execute_reply.started":"2023-03-27T19:46:05.030467Z","shell.execute_reply":"2023-03-27T19:46:05.389469Z"}}},{"cell_type":"code","source":"import random\n\n# Generating random index to test\ntest_index = random.randint(0,len(X))","metadata":{"execution":{"iopub.status.busy":"2023-03-29T01:00:07.010122Z","iopub.execute_input":"2023-03-29T01:00:07.010658Z","iopub.status.idle":"2023-03-29T01:00:07.026446Z","shell.execute_reply.started":"2023-03-29T01:00:07.010594Z","shell.execute_reply":"2023-03-29T01:00:07.023720Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Should return spectrogram and numpy array of classes\nX_function_test, y_function_test = preprocess(X[test_index], y[test_index])","metadata":{"execution":{"iopub.status.busy":"2023-03-29T01:00:07.029062Z","iopub.execute_input":"2023-03-29T01:00:07.029632Z","iopub.status.idle":"2023-03-29T01:00:07.261364Z","shell.execute_reply.started":"2023-03-29T01:00:07.029584Z","shell.execute_reply":"2023-03-29T01:00:07.258328Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Verifying datatypes and shapes\nprint(f'Spectrogram shape: {X_function_test.shape}')\nprint(f'Spectrogram type: {type(X_function_test)}')\nprint(f'Classes shape: {y_function_test.shape}')\nprint(f'Classes type: {type(y_function_test)}')","metadata":{"execution":{"iopub.status.busy":"2023-03-29T01:00:07.264137Z","iopub.execute_input":"2023-03-29T01:00:07.265010Z","iopub.status.idle":"2023-03-29T01:00:07.271870Z","shell.execute_reply.started":"2023-03-29T01:00:07.264956Z","shell.execute_reply":"2023-03-29T01:00:07.270512Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### 2.6 Plotting Spectrogram","metadata":{}},{"cell_type":"code","source":"plt.figure(figsize=(10, 4))\nplt.imshow(tf.transpose(X_function_test)[0], aspect='auto', origin='lower', cmap='viridis')\nplt.colorbar()\nplt.xlabel('Time')\nplt.ylabel('Frequency')\nplt.title('Bird Call Spectrogram (Downsampled to 16 kHz)')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-03-29T01:00:07.273993Z","iopub.execute_input":"2023-03-29T01:00:07.275128Z","iopub.status.idle":"2023-03-29T01:00:07.639013Z","shell.execute_reply.started":"2023-03-29T01:00:07.275071Z","shell.execute_reply":"2023-03-29T01:00:07.637910Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 3. Data Transforming\nNow that there is a defined function that can preprocess each tuple of (input, output) data, the next step is to create a dataset using TensorFlow and define the data pipeline.","metadata":{}},{"cell_type":"markdown","source":"### 3.1 Creating dataset","metadata":{}},{"cell_type":"code","source":"# Creating dataset from 'train_metadata.csv' data\ndata = tf.data.Dataset.from_tensor_slices((X, y))","metadata":{"execution":{"iopub.status.busy":"2023-03-29T01:00:07.640391Z","iopub.execute_input":"2023-03-29T01:00:07.641614Z","iopub.status.idle":"2023-03-29T01:00:07.667847Z","shell.execute_reply.started":"2023-03-29T01:00:07.641499Z","shell.execute_reply":"2023-03-29T01:00:07.664681Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### 3.2 Creating data pipeline","metadata":{}},{"cell_type":"code","source":"# Full data pipeline\ndata = data.map(preprocess, num_parallel_calls=tf.data.AUTOTUNE)\ndata = data.cache()\ndata = data.shuffle(buffer_size=100)\ndata = data.batch(batch_size=16)\ndata = data.prefetch(8)","metadata":{"execution":{"iopub.status.busy":"2023-03-29T01:00:07.670682Z","iopub.execute_input":"2023-03-29T01:00:07.671286Z","iopub.status.idle":"2023-03-29T01:00:07.938978Z","shell.execute_reply.started":"2023-03-29T01:00:07.671200Z","shell.execute_reply":"2023-03-29T01:00:07.936410Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### 3.3 Training and Testing Partitions","metadata":{}},{"cell_type":"code","source":"# Calculating batch numbers for train and test\nbatches = len(data)\nprint(batches)\nprint(batches * (0.8))\nprint(batches - (batches * (0.8)))","metadata":{"execution":{"iopub.status.busy":"2023-03-29T01:00:07.941046Z","iopub.execute_input":"2023-03-29T01:00:07.941476Z","iopub.status.idle":"2023-03-29T01:00:07.951347Z","shell.execute_reply.started":"2023-03-29T01:00:07.941436Z","shell.execute_reply":"2023-03-29T01:00:07.948813Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# 847 Training batches\n# 212 Test batches\ntrain_1 = data.take(200)\ntrain_2 = data.skip(200).take(200)\ntrain_3 = data.skip(400).take(200)\ntrain_4 = data.skip(600).take(247)\ntest = data.skip(847).take(212)","metadata":{"execution":{"iopub.status.busy":"2023-03-29T01:00:07.953383Z","iopub.execute_input":"2023-03-29T01:00:07.953958Z","iopub.status.idle":"2023-03-29T01:00:07.972766Z","shell.execute_reply.started":"2023-03-29T01:00:07.953906Z","shell.execute_reply":"2023-03-29T01:00:07.970254Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### 3.4 Verifying Data Generator","metadata":{}},{"cell_type":"code","source":"# Should output ([batch_size], [element dimension]) for both input and output\nspectrograms, labels = train_1.as_numpy_iterator().next()","metadata":{"execution":{"iopub.status.busy":"2023-03-29T01:00:07.976012Z","iopub.execute_input":"2023-03-29T01:00:07.976643Z","iopub.status.idle":"2023-03-29T01:00:11.514634Z","shell.execute_reply.started":"2023-03-29T01:00:07.976548Z","shell.execute_reply":"2023-03-29T01:00:11.512793Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# 16 examples of spectrograms\nspectrograms.shape","metadata":{"execution":{"iopub.status.busy":"2023-03-29T01:00:11.516032Z","iopub.execute_input":"2023-03-29T01:00:11.516331Z","iopub.status.idle":"2023-03-29T01:00:11.525604Z","shell.execute_reply.started":"2023-03-29T01:00:11.516302Z","shell.execute_reply":"2023-03-29T01:00:11.523450Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# 16 examples of species label data\nlabels.shape","metadata":{"execution":{"iopub.status.busy":"2023-03-29T01:00:11.528349Z","iopub.execute_input":"2023-03-29T01:00:11.530199Z","iopub.status.idle":"2023-03-29T01:00:11.543758Z","shell.execute_reply.started":"2023-03-29T01:00:11.530066Z","shell.execute_reply":"2023-03-29T01:00:11.542486Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 4. CNN Model Development","metadata":{}},{"cell_type":"code","source":"# TensorFlow CNN Model Dependencies\nfrom tensorflow.keras.models import Sequential\nfrom tensorflow.keras.layers import Conv2D, Flatten, Dense, MaxPooling2D","metadata":{"execution":{"iopub.status.busy":"2023-03-29T01:00:11.545499Z","iopub.execute_input":"2023-03-29T01:00:11.545868Z","iopub.status.idle":"2023-03-29T01:00:11.561776Z","shell.execute_reply.started":"2023-03-29T01:00:11.545833Z","shell.execute_reply":"2023-03-29T01:00:11.559756Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### 4.1 Define Model","metadata":{}},{"cell_type":"code","source":"model = Sequential()\n\n# Input layer (takes in spectrogram)\nmodel.add(Conv2D(16, (3, 3), activation='relu', input_shape=(1491, 257, 1)))\nmodel.add(MaxPooling2D(2, 2))\n\nmodel.add(Conv2D(16, (3, 3), activation='relu'))\nmodel.add(MaxPooling2D(2, 2))\n\n# Flattening layer -> Dense\nmodel.add(Flatten())\nmodel.add(Dense(264, activation='relu'))\n\n# Output layer (probabilities)\nmodel.add(Dense(264, activation='softmax'))\n\nmodel.compile('Adam', loss='CategoricalCrossentropy',\n              metrics=[tf.keras.metrics.CategoricalAccuracy()])","metadata":{"execution":{"iopub.status.busy":"2023-03-29T01:00:11.564437Z","iopub.execute_input":"2023-03-29T01:00:11.564928Z","iopub.status.idle":"2023-03-29T01:00:12.554312Z","shell.execute_reply.started":"2023-03-29T01:00:11.564884Z","shell.execute_reply":"2023-03-29T01:00:12.552946Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### 4.2 View Summary","metadata":{}},{"cell_type":"code","source":"model.summary()","metadata":{"execution":{"iopub.status.busy":"2023-03-29T01:00:12.555802Z","iopub.execute_input":"2023-03-29T01:00:12.556174Z","iopub.status.idle":"2023-03-29T01:00:12.599687Z","shell.execute_reply.started":"2023-03-29T01:00:12.556134Z","shell.execute_reply":"2023-03-29T01:00:12.598071Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### 4.3 Fit Model, View KPIs and Plot","metadata":{}},{"cell_type":"code","source":"from tensorflow.keras.callbacks import EarlyStopping\n\n# Stops model if it is not steadily dropping loss val_loss function return\nearly_stopping = EarlyStopping(monitor='val_loss', patience=3, restore_best_weights=True)","metadata":{"execution":{"iopub.status.busy":"2023-03-29T01:00:12.601798Z","iopub.execute_input":"2023-03-29T01:00:12.602191Z","iopub.status.idle":"2023-03-29T01:00:12.609403Z","shell.execute_reply.started":"2023-03-29T01:00:12.602146Z","shell.execute_reply":"2023-03-29T01:00:12.608169Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Fitting model to train_1\nhistory_train_1 = model.fit(train_1, epochs=100, \n                            validation_data=test, callbacks=[early_stopping])","metadata":{"execution":{"iopub.status.busy":"2023-03-29T01:00:12.610709Z","iopub.execute_input":"2023-03-29T01:00:12.611032Z"},"trusted":true},"execution_count":null,"outputs":[]}]}