{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"In this notebook, we'll use a pre-trained machine learning model to generate a submission to the [BirdClef2023 competition](https://www.kaggle.com/c/birdclef-2023).  The goal of the competition is to identify Eastern African bird species by sound.","metadata":{}},{"cell_type":"markdown","source":"## Step 1: Imports","metadata":{}},{"cell_type":"code","source":"import tensorflow as tf\nimport tensorflow_hub as hub # Dépôt de modèles de machine learning entraînés, prêts à être optimisés et déployés n'importe où\nimport tensorflow_io as tfio # Ensemble de systèmes et de formats de fichiers qui ne sont pas compatibles en mode natif avec TensorFlow\n\nimport pandas as pd\nimport numpy as np\nimport librosa # Analyse musicale et audio - Permet de créer des systèmes de recherche d'informations musicales\nimport glob # Génère des paths en utilisant les règles de regex du UNIX shell\n\nimport csv\nimport io # Librairie par défaut pour gérer les flux et fichiers\n\nfrom IPython.display import Audio # Permet de lire des fichiers audio","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2023-03-28T14:42:02.085004Z","iopub.execute_input":"2023-03-28T14:42:02.085459Z","iopub.status.idle":"2023-03-28T14:42:12.644289Z","shell.execute_reply.started":"2023-03-28T14:42:02.085409Z","shell.execute_reply":"2023-03-28T14:42:12.642658Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Step 2: Explore the training data\n\nWe'll start by loading a couple of training examples and using the IPython.display.Audio module to play them!","metadata":{}},{"cell_type":"code","source":"# Load a sample audio files from two different species\naudio_abe, sr_abe = librosa.load(\"/kaggle/input/birdclef-2023/train_audio/abethr1/XC128013.ogg\")\naudio_abh, sr_abh = librosa.load(\"/kaggle/input/birdclef-2023/train_audio/abhori1/XC127317.ogg\")\n\n# audio=audio times series, multi-channel is supported - sr = sampling rate (taux d'échantillonage)","metadata":{"execution":{"iopub.status.busy":"2023-03-28T14:42:12.646789Z","iopub.execute_input":"2023-03-28T14:42:12.647594Z","iopub.status.idle":"2023-03-28T14:42:25.762665Z","shell.execute_reply.started":"2023-03-28T14:42:12.647545Z","shell.execute_reply":"2023-03-28T14:42:25.761171Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The sample rate is a measure of how often the amplitude of an analog signal (such as an audio waveform) is measured and recorded as a digital signal. In other words, it represents the number of samples per second that are taken from the analog signal to create a digital representation of the signal.\n\nIn the provided code, the sample rate is specified as 32,000 Hz (or 32 kHz) in several places. This means that the audio waveform is being sampled 32,000 times per second to create a digital representation of the waveform. This is a common sample rate used in audio processing and is often sufficient for accurately capturing the frequencies present in many types of audio signals.\n\nIt's worth noting that the sample rate can have an impact on the quality of the digital representation of the audio waveform. A higher sample rate can capture more detail in the waveform and result in a more accurate representation of the original signal, but also requires more storage space and processing power. Conversely, a lower sample rate can result in a lower-quality representation of the signal, but requires less storage and processing power.","metadata":{}},{"cell_type":"code","source":"# Play the audio\nAudio(data=audio_abe, rate=sr_abe)","metadata":{"execution":{"iopub.status.busy":"2023-03-28T14:42:25.764644Z","iopub.execute_input":"2023-03-28T14:42:25.765734Z","iopub.status.idle":"2023-03-28T14:42:25.848483Z","shell.execute_reply.started":"2023-03-28T14:42:25.765673Z","shell.execute_reply":"2023-03-28T14:42:25.846193Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Play the audio\nAudio(data=audio_abh, rate=sr_abh)","metadata":{"execution":{"iopub.status.busy":"2023-03-28T14:42:25.852754Z","iopub.execute_input":"2023-03-28T14:42:25.853891Z","iopub.status.idle":"2023-03-28T14:42:25.926500Z","shell.execute_reply.started":"2023-03-28T14:42:25.853835Z","shell.execute_reply":"2023-03-28T14:42:25.925262Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Step 3: Match the model's output with the bird species in the competition\n\nThe competition includes 264 classes of birds, 261 of which exist in this model. We'll set up a way to map the model's output logits to our competition.","metadata":{}},{"cell_type":"code","source":"model = hub.load('https://kaggle.com/models/google/bird-vocalization-classifier/frameworks/tensorFlow2/variations/bird-vocalization-classifier/versions/1')\nlabels_path = hub.resolve('https://kaggle.com/models/google/bird-vocalization-classifier/frameworks/tensorFlow2/variations/bird-vocalization-classifier/versions/1') + \"/assets/label.csv\"","metadata":{"execution":{"iopub.status.busy":"2023-03-28T14:42:25.928404Z","iopub.execute_input":"2023-03-28T14:42:25.929040Z","iopub.status.idle":"2023-03-28T14:42:34.288114Z","shell.execute_reply.started":"2023-03-28T14:42:25.928985Z","shell.execute_reply":"2023-03-28T14:42:34.286780Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"### Input: 5 seconds of silence as mono 32 kHz waveform samples.\nwaveform = np.zeros(shape=(1,160000), dtype=np.float32)\n\n### Run the model, check the output.\nlogits, embeddings = model.infer_tf(waveform)\nprint(logits.shape)\nprint(embeddings.shape)\n\n\"\"\"\nIn machine learning, an embedding is a representation of a high-dimensional vector or sequence of data points in a lower-dimensional space. The goal of an embedding is to capture the essential characteristics of the data while reducing its dimensionality. Embeddings are widely used in natural language processing, computer vision, and audio processing, among other fields.\nIn the case of audio processing, embeddings are commonly used to represent audio signals in a lower-dimensional space, where they can be more easily compared, classified, or visualized. The embeddings learned by the bird vocalization classifier model represent the audio waveform in a high-dimensional space, where similar audio signals are grouped together, and dissimilar ones are far apart.\nHere's an example of how embeddings can be used in machine learning:\nSuppose you have a dataset of audio recordings of different bird species, and you want to classify each recording based on its species. You can use a pre-trained model such as the bird vocalization classifier to extract the embeddings of each audio recording, and then train a classifier such as a neural network on these embeddings to predict the bird species.\nDuring training, the classifier learns to map the embeddings of each audio recording to its corresponding bird species. Once trained, the classifier can predict the species of new audio recordings based on their embeddings.\nIn summary, embeddings are a powerful tool in machine learning for representing complex data such as audio signals in a lower-dimensional space, where they can be more easily analyzed, compared, or visualized.\n\"\"\"","metadata":{"execution":{"iopub.status.busy":"2023-03-28T14:42:34.292054Z","iopub.execute_input":"2023-03-28T14:42:34.293120Z","iopub.status.idle":"2023-03-28T14:42:45.791830Z","shell.execute_reply.started":"2023-03-28T14:42:34.293060Z","shell.execute_reply":"2023-03-28T14:42:45.790528Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(labels_path)\ndf = pd.read_csv(\"/kaggle/input/bird-vocalization-classifier/tensorflow2/bird-vocalization-classifier/1/assets/label.csv\")\ndf.head","metadata":{"execution":{"iopub.status.busy":"2023-03-28T14:42:45.793057Z","iopub.execute_input":"2023-03-28T14:42:45.793383Z","iopub.status.idle":"2023-03-28T14:42:45.841056Z","shell.execute_reply.started":"2023-03-28T14:42:45.793349Z","shell.execute_reply":"2023-03-28T14:42:45.840110Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Find the name of the class with the top score when mean-aggregated across frames.\ndef class_names_from_csv(class_map_csv_text):\n    \"\"\"Returns list of class names corresponding to score vector.\"\"\"\n    with open(labels_path) as csv_file:\n        csv_reader = csv.reader(csv_file, delimiter=',')\n        class_names = [mid for mid, desc in csv_reader] # (mid, desc) = (nom, commentaire)\n        return class_names[1:]\n\n## note that the bird classifier classifies a much larger set of birds than the\n## competition, so we need to load the model's set of class names or else our \n## indices will be off.\nclasses = class_names_from_csv(labels_path) # Liste des noms du modèle TF","metadata":{"execution":{"iopub.status.busy":"2023-03-28T14:42:45.842381Z","iopub.execute_input":"2023-03-28T14:42:45.843317Z","iopub.status.idle":"2023-03-28T14:42:45.855352Z","shell.execute_reply.started":"2023-03-28T14:42:45.843275Z","shell.execute_reply":"2023-03-28T14:42:45.854122Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_metadata = pd.read_csv(\"/kaggle/input/birdclef-2023/train_metadata.csv\")\ntrain_metadata.head()","metadata":{"execution":{"iopub.status.busy":"2023-03-28T15:56:44.049266Z","iopub.execute_input":"2023-03-28T15:56:44.050609Z","iopub.status.idle":"2023-03-28T15:56:44.134301Z","shell.execute_reply.started":"2023-03-28T15:56:44.050533Z","shell.execute_reply":"2023-03-28T15:56:44.133118Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"competition_classes = sorted(train_metadata.primary_label.unique())\n\nforced_defaults = 0\ncompetition_class_map = [] # En 0, l'indice TF du 0ième élement de notre table...\nfor c in competition_classes: # Parcourt les labels de notre table\n    try:\n        i = classes.index(c) # Indice du label dans la liste des noms obtenue avec le modèle TF\n        competition_class_map.append(i)\n    except:\n        competition_class_map.append(0) # Si on ne trouve pas le label dans la liste TF, on ajoute 0\n        forced_defaults += 1 # On compte le nombre de missing labels \n        \n## this is the count of classes not supported by our pretrained model\n## you could choose to simply not predict these, set a default as above,\n## or create your own model using the pretrained model as a base.\nforced_defaults\nprint(len(competition_class_map)) # 264 labels uniques!!!\nprint(competition_class_map)","metadata":{"execution":{"iopub.status.busy":"2023-03-28T16:04:09.328991Z","iopub.execute_input":"2023-03-28T16:04:09.329451Z","iopub.status.idle":"2023-03-28T16:04:09.373155Z","shell.execute_reply.started":"2023-03-28T16:04:09.329402Z","shell.execute_reply":"2023-03-28T16:04:09.371874Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Step 4: Preprocess the data\n\nThe following functions are one way to load the audio provided and break it up into the five-second samples with a sample rate of 32,000 required by the competition.","metadata":{}},{"cell_type":"code","source":"def frame_audio(\n      audio_array: np.ndarray,\n      window_size_s: float = 5.0,\n      hop_size_s: float = 5.0, # On découpe vraiment en segments de 5 secondes!\n      sample_rate = 32000,\n      ) -> np.ndarray:\n    \n    \"\"\"Helper function for framing audio for inference.\"\"\"\n    \"\"\" using tf.signal \"\"\"\n    if window_size_s is None or window_size_s < 0:\n        return audio_array[np.newaxis, :]\n    frame_length = int(window_size_s * sample_rate)\n    hop_length = int(hop_size_s * sample_rate)\n    framed_audio = tf.signal.frame(audio_array, frame_length, hop_length, pad_end=True) # pad_value=0 par défaut\n    # Lire https://www.tensorflow.org/api_docs/python/tf/signal/frame\n    return framed_audio\n\ndef ensure_sample_rate(waveform, original_sample_rate,\n                       desired_sample_rate=32000):\n    \"\"\"Resample waveform if required.\"\"\"\n    if original_sample_rate != desired_sample_rate:\n        waveform = tfio.audio.resample(waveform, original_sample_rate, desired_sample_rate)\n    return desired_sample_rate, waveform","metadata":{"execution":{"iopub.status.busy":"2023-03-28T14:42:46.059148Z","iopub.execute_input":"2023-03-28T14:42:46.059614Z","iopub.status.idle":"2023-03-28T14:42:46.070996Z","shell.execute_reply.started":"2023-03-28T14:42:46.059566Z","shell.execute_reply":"2023-03-28T14:42:46.069166Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"These are two helper functions used to prepare audio for inference in the competition:\n\n    frame_audio: This function takes an audio array (audio_array) as input and frames the audio into segments of a fixed length (specified by window_size_s) with a specified overlap between adjacent segments (specified by hop_size_s). The function uses the frame function from the tf.signal module in TensorFlow to perform the framing operation. The resulting framed audio is returned as a NumPy array.\n\n    ensure_sample_rate: This function takes a waveform (which may be a framed audio segment output by frame_audio) and its original sample rate (original_sample_rate) as input. If the original sample rate is not equal to the desired sample rate of 32,000 Hz, the function uses the tfio.audio.resample function from the TensorFlow I/O module to resample the waveform to the desired sample rate. The function returns the desired sample rate and the resampled waveform.\n\nThese functions are useful for preparing audio for inference with the provided bird vocalization classifier model, which requires audio to be framed into segments of a fixed length with a specific sample rate. The functions can be used to preprocess audio data prior to feeding it into the model for classification.","metadata":{}},{"cell_type":"markdown","source":"Below we load one training sample - use the Audio function to listen to the samples inside the notebook!","metadata":{}},{"cell_type":"code","source":"audio, sample_rate = librosa.load(\"/kaggle/input/birdclef-2023/train_audio/afghor1/XC156639.ogg\")\nsample_rate, wav_data = ensure_sample_rate(audio, sample_rate)\nAudio(wav_data, rate=sample_rate)","metadata":{"execution":{"iopub.status.busy":"2023-03-28T14:42:46.072712Z","iopub.execute_input":"2023-03-28T14:42:46.073248Z","iopub.status.idle":"2023-03-28T14:42:46.825806Z","shell.execute_reply.started":"2023-03-28T14:42:46.073196Z","shell.execute_reply":"2023-03-28T14:42:46.822301Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Step 5: Make predictions\n\nEach test sample is cut into 5-second chunks. We use the pretrained model to return probabilities for all 10k birds included in the model, then pull out the classes used in this competition to create a final submission row. Note that we are NOT doing anything special to handle the 3 missing classes; those will need fine-tuning / transfer learning, which will be handled in a separate notebook.","metadata":{}},{"cell_type":"code","source":"fixed_tm = frame_audio(wav_data)\nprint(fixed_tm.shape)\nlogits, embeddings = model.infer_tf(fixed_tm[:1])\nprobabilities = tf.nn.softmax(logits) # On se ramène à des probabilités\nargmax = np.argmax(probabilities) # Indice de la proba la plus élevée\nprint(f\"The audio is from the class {classes[argmax]} (element:{argmax} in the label.csv file), with probability of {probabilities[0][argmax]}\")\n\n# classes obtenu avec modèle TF","metadata":{"execution":{"iopub.status.busy":"2023-03-28T14:50:13.965227Z","iopub.execute_input":"2023-03-28T14:50:13.965641Z","iopub.status.idle":"2023-03-28T14:50:16.366029Z","shell.execute_reply.started":"2023-03-28T14:50:13.965603Z","shell.execute_reply":"2023-03-28T14:50:16.364532Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"![https://i.stack.imgur.com/zkMBy.png](https://i.stack.imgur.com/zkMBy.png)","metadata":{}},{"cell_type":"code","source":"def predict_for_sample(filename, sample_submission, frame_limit_secs=None): # filename is a path\n    file_id = filename.split(\".ogg\")[0].split(\"/\")[-1]\n    \n    audio, sample_rate = librosa.load(filename) # sr par défaut 22050\n    sample_rate, wav_data = ensure_sample_rate(audio, sample_rate)\n    \n    fixed_tm = frame_audio(wav_data)\n    \n    frame = 5\n    all_logits, all_embeddings = model.infer_tf(fixed_tm[:1]) # Toute la première ligne qui contient le premier extrait de 5 secondes\n    for window in fixed_tm[1:]:\n        if frame_limit_secs and frame > frame_limit_secs: # if nombre -> True, vérifie que la frame length est au dessus d'une taille max\n            continue\n        \n        logits, embeddings = model.infer_tf(window[np.newaxis, :])\n        all_logits = np.concatenate([all_logits, logits], axis=0) # https://numpy.org/doc/stable/reference/generated/numpy.concatenate.html\n        print(all_logits)\n        frame += 5\n    \n    frame = 5\n    all_probabilities = []\n    for frame_logits in all_logits:\n        probabilities = tf.nn.softmax(frame_logits).numpy()\n        \n        ## set the appropriate row in the sample submission\n        sample_submission.loc[sample_submission.row_id == file_id + \"_\" + str(frame), competition_classes] = probabilities[competition_class_map]\n        # sample_submission[True, False, True, ... | toutes les colonnes en sorted] = \n        frame += 5","metadata":{"execution":{"iopub.status.busy":"2023-03-28T15:12:45.232316Z","iopub.execute_input":"2023-03-28T15:12:45.232747Z","iopub.status.idle":"2023-03-28T15:12:45.244419Z","shell.execute_reply.started":"2023-03-28T15:12:45.232712Z","shell.execute_reply":"2023-03-28T15:12:45.243016Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Here is a brief explanation of each step in the code:\n\n    file_id = filename.split(\".ogg\")[0].split(\"/\")[-1] - This extracts the unique identifier for the audio file from its filename. The filename is assumed to be in the format file_id.ogg.\n\n    audio, sample_rate = librosa.load(filename) - This loads the audio file using the librosa library and returns the audio data as a NumPy array and its sample rate.\n\n    sample_rate, wav_data = ensure_sample_rate(audio, sample_rate) - This resamples the audio data to a fixed sample rate of 32kHz using the ensure_sample_rate function.\n\n    fixed_tm = frame_audio(wav_data) - This breaks up the resampled audio data into fixed-length frames of 5 seconds each using the frame_audio function.\n\n    all_logits, all_embeddings = model.infer_tf(fixed_tm[:1]) - This uses the pre-trained TensorFlow model to make predictions on the first 5-second frame of the audio data and returns the logits (i.e., unnormalized class scores) and embeddings (i.e., lower-dimensional feature representations) for each segment.\n\n    for window in fixed_tm[1:]: - This loops over the remaining frames of the audio data and makes predictions on each frame using the same pre-trained model.\n\n    if frame_limit_secs and frame > frame_limit_secs: - This checks if the frame length exceeds a specified time limit and skips over it if it does.\n\n    logits, embeddings = model.infer_tf(window[np.newaxis, :]) - This uses the pre-trained model to make predictions on the current frame and returns the logits and embeddings for each class.\n\n    all_logits = np.concatenate([all_logits, logits], axis=0) - This appends the logits for the current frame to the array of logits for all frames.\n\n    probabilities = tf.nn.softmax(frame_logits).numpy() - This computes the softmax probabilities of the class logits for the current frame.\n\n    sample_submission.loc[sample_submission.row_id == file_id + \"_\" + str(frame), competition_classes] = probabilities[competition_class_map] - This updates the corresponding row in the sample submission DataFrame with the predicted class probabilities.\n\n    frame += 5 - This increments the frame index by 5 seconds to process the next frame in the audio file.\n\n","metadata":{}},{"cell_type":"markdown","source":"## Step 6: Generate a submission\n\nNow we process all of the test samples as discussed above, creating output rows, and saving them in the provided `sample_submission.csv`. Finally, we save these rows to our final output file: `submission.csv`. This is the file that gets submitted and scored when you submit the notebook.","metadata":{}},{"cell_type":"code","source":"test_samples = list(glob.glob(\"/kaggle/input/birdclef-2023/test_soundscapes/*.ogg\"))\ntest_samples","metadata":{"execution":{"iopub.status.busy":"2023-03-28T15:33:57.290182Z","iopub.execute_input":"2023-03-28T15:33:57.290634Z","iopub.status.idle":"2023-03-28T15:33:57.302635Z","shell.execute_reply.started":"2023-03-28T15:33:57.290594Z","shell.execute_reply":"2023-03-28T15:33:57.301195Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sample_sub = pd.read_csv(\"/kaggle/input/birdclef-2023/sample_submission.csv\")\nprint(sample_sub.shape)\nprint(sample_sub[competition_classes].shape) \nprint(len(competition_classes))\n\nsample_sub[competition_classes] = sample_sub[competition_classes].astype(np.float32)\nsample_sub.head()","metadata":{"execution":{"iopub.status.busy":"2023-03-28T16:02:36.313800Z","iopub.execute_input":"2023-03-28T16:02:36.314253Z","iopub.status.idle":"2023-03-28T16:02:36.420661Z","shell.execute_reply.started":"2023-03-28T16:02:36.314206Z","shell.execute_reply":"2023-03-28T16:02:36.419707Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"frame_limit_secs = 15 if sample_sub.shape[0] == 3 else None\nfor sample_filename in test_samples:\n    predict_for_sample(sample_filename, sample_sub, frame_limit_secs=15)","metadata":{"execution":{"iopub.status.busy":"2023-03-28T15:49:00.980837Z","iopub.execute_input":"2023-03-28T15:49:00.981869Z","iopub.status.idle":"2023-03-28T15:49:13.614599Z","shell.execute_reply.started":"2023-03-28T15:49:00.981820Z","shell.execute_reply":"2023-03-28T15:49:13.613496Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sample_sub","metadata":{"execution":{"iopub.status.busy":"2023-03-28T15:49:18.520735Z","iopub.execute_input":"2023-03-28T15:49:18.521798Z","iopub.status.idle":"2023-03-28T15:49:18.553124Z","shell.execute_reply.started":"2023-03-28T15:49:18.521750Z","shell.execute_reply":"2023-03-28T15:49:18.551952Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sample_sub.to_csv(\"submission.csv\", index=False)","metadata":{"execution":{"iopub.status.busy":"2023-03-28T15:49:20.108824Z","iopub.execute_input":"2023-03-28T15:49:20.109248Z","iopub.status.idle":"2023-03-28T15:49:20.130882Z","shell.execute_reply.started":"2023-03-28T15:49:20.109211Z","shell.execute_reply":"2023-03-28T15:49:20.129479Z"},"trusted":true},"execution_count":null,"outputs":[]}]}