{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"**Part 0: Exploratory Data Analysis**: https://www.kaggle.com/code/virajkadam/birdclef2022-data-and-audio-eda\n\n\n**Part 2: Training a No-call Classifier**:https://www.kaggle.com/code/virajkadam/birdclef22-p2-no-call-classifier","metadata":{}},{"cell_type":"markdown","source":"# Introduction\n\n<h3 style= 'color:green'> The objective the competition is to classify the birds using the bird audio recordings. </h3>\n    \n\n**Extracting Spectrograms**\n\nFrom the blog here : https://www.macaulaylibrary.org/2021/07/19/from-sound-to-images-part-1-a-deep-dive-on-spectrogram-creation/\n\nWhat even is a Spectrogram?\n\n    A spectrogram tracks the sound frequencies (vertical axis) which appear in the waveform, as a function of time (horizontal axis). Brighter colors correspond to louder sounds.\n\nWhy use spectrograms? \n\n    Rather than working with waveforms directly, we have the option of representing our sound as an image either as a spectrogram or some other image representation. This approach has several advantages.\n    First, image representations are indispensable tools for bird sound ID experts who are trying to identify species in a recording.\n    Second, by using image representations, we have the option of using well-understood computer vision model architectures like RestNets along with pretrained weights from Imagenet.\n    \n**The NO Call classifier**    \n\n    As a part of training , we will convert the audio recordings into 10-second spectograms. These individual spectograms are not guaranteed to have a bird call in them, and hence it is important to filter out the empty/noisy portions of the audio for better training. A lot of the test recordings also do not have a bird call in them. \n\n\n    Hence we will train a classifier to filter out no-call samples here : https://www.kaggle.com/code/virajkadam/birdclef22-p2-no-call-classifier/edit","metadata":{}},{"cell_type":"markdown","source":"**This is the step 1 of the training pipeline. Link to the next notebook:**","metadata":{}},{"cell_type":"markdown","source":"# Creating Spectogram : \n\nfrom : https://www.macaulaylibrary.org/2021/07/19/from-sound-to-images-part-1-a-deep-dive-on-spectrogram-creation/\n\nThere are a number of choices one can make when constructing a spectrogram. Among them are the following:\n\n• Clip length: How many seconds of audio should a spectrogram represent?\n\n    > Shorter clips often eliminated important context from the soundscape,while longer clips are more expensive to train.\n\n• STFT window length: Represents a tradeoff between a spectrogram’s level of resolution in the time domain (short window length) and resolution in the frequency domain (long window length). \n\n    > 256 or 512 samples \n\n• STFT hop size: A shorter hop size leads to higher resolution in the time domain, but results in larger inputs to the model. In turn, larger inputs may slow down model training and inference.\n    \n    > 64 or 128 samples\n    \n• Mel scaling: In a frequency spectrogram, the vertical distance that represents an octave is not constant. As a result, it may be difficult for convolutional filters to learn to recognize harmonies, overtones, and repeated harmonic patterns. Mel scaling rescales the frequency axis, so that fixed differences in musical pitch (e.g. an octave or a fifth) correspond to fixed vertical distances. One possible downside to using mel scaling is that high frequency sounds will become compressed at the top of the spectrogram, and therefore might be harder to distinguish.\n    \n    > On or off\n\n• Image rescaling: Choices of the parameters above affect the spatial dimensions of the resulting spectrogram. To make meaningful comparisons, we chose a set of image dimensions to rescale our spectrograms to.\n\n    >rescale to [128, 512] or [96, 512] (for hop size 128), or [128, 1024] (for hop size 64)","metadata":{}},{"cell_type":"markdown","source":"# Imports","metadata":{}},{"cell_type":"code","source":"!pip install audiomentations -q\n!pip install pqdm -q\n\nimport os\nimport pandas as pd\nimport numpy as np\nimport json\n\nfrom pathlib import Path\nimport pqdm\nfrom tqdm import tqdm\n\n\nimport matplotlib.pyplot as plt\nfrom PIL import Image\n\n\nimport warnings\nwarnings.filterwarnings(action='ignore')\n\n#audio\nimport librosa\nfrom IPython.display import Audio\n#audio augmentations'\nfrom audiomentations import Compose,AddGaussianSNR,Shift,TimeStretch,TimeMask,FrequencyMask,PolarityInversion","metadata":{"execution":{"iopub.status.busy":"2022-05-02T08:22:47.861904Z","iopub.execute_input":"2022-05-02T08:22:47.862841Z","iopub.status.idle":"2022-05-02T08:23:09.665560Z","shell.execute_reply.started":"2022-05-02T08:22:47.862736Z","shell.execute_reply":"2022-05-02T08:23:09.664507Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Config","metadata":{}},{"cell_type":"code","source":"# Global vars\nRANDOM_SEED = 7\nSAMPLE_RATE = 22000\nSIGNAL_LENGTH = 10 # seconds\nSPEC_SHAPE = (128, 512) # height x width\nFMIN = 500\nFMAX = 12500\nhop_length = int(SIGNAL_LENGTH * SAMPLE_RATE / (SPEC_SHAPE[1] - 1))","metadata":{"execution":{"iopub.status.busy":"2022-05-02T08:23:09.668046Z","iopub.execute_input":"2022-05-02T08:23:09.668839Z","iopub.status.idle":"2022-05-02T08:23:09.674572Z","shell.execute_reply.started":"2022-05-02T08:23:09.668779Z","shell.execute_reply":"2022-05-02T08:23:09.673144Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Loading data**","metadata":{}},{"cell_type":"code","source":"def json_to_pd(file):\n    '''read and conv json file to pd row'''\n    with open(file) as f:\n        json_data = pd.json_normalize(json.loads(f.read()))\n        \n    return json_data\n\n","metadata":{"execution":{"iopub.status.busy":"2022-05-02T08:23:09.675927Z","iopub.execute_input":"2022-05-02T08:23:09.676215Z","iopub.status.idle":"2022-05-02T08:23:09.689134Z","shell.execute_reply.started":"2022-05-02T08:23:09.676185Z","shell.execute_reply":"2022-05-02T08:23:09.688346Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#BIRDCLEF 22 DATA \n\n\ntrain_audio = '../input/birdclef-2022/train_audio'\ntrain_metadata = pd.read_csv('../input/birdclef-2022/train_metadata.csv')\n\n#taking samples with rating > 2.5 \ntrain_metadata=train_metadata.query('rating>=2.5')\n\n\n#make directories to save spectograms\nTrain_Spectrograms = './Train_Spectrograms'\nFreefield_Spectograms = './Freefield_Spectrograms'\n\n!mkdir $Train_Spectrograms\n!mkdir $Freefield_Spectograms","metadata":{"execution":{"iopub.status.busy":"2022-05-02T08:23:09.691074Z","iopub.execute_input":"2022-05-02T08:23:09.691510Z","iopub.status.idle":"2022-05-02T08:23:11.383903Z","shell.execute_reply.started":"2022-05-02T08:23:09.691480Z","shell.execute_reply":"2022-05-02T08:23:11.382670Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_metadata.shape","metadata":{"execution":{"iopub.status.busy":"2022-05-02T08:23:11.386232Z","iopub.execute_input":"2022-05-02T08:23:11.387184Z","iopub.status.idle":"2022-05-02T08:23:11.396950Z","shell.execute_reply.started":"2022-05-02T08:23:11.387125Z","shell.execute_reply":"2022-05-02T08:23:11.396264Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#add a filepath to each audio \n\ntrain_metadata['filepath'] = train_audio + '/' + train_metadata['filename']\ntrain_metadata.head()","metadata":{"execution":{"iopub.status.busy":"2022-05-02T08:23:11.398276Z","iopub.execute_input":"2022-05-02T08:23:11.398692Z","iopub.status.idle":"2022-05-02T08:23:11.429904Z","shell.execute_reply.started":"2022-05-02T08:23:11.398664Z","shell.execute_reply":"2022-05-02T08:23:11.429318Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Freefield Data**\n\nABOUT\n\nThis dataset contains 7690 10-second audio files in a standardised format, extracted from contributions on the Freesound archive which were labelled with the \"field-recording\" tag. Note that the original tagging (as well as the audio submission) is crowdsourced, so the dataset is not guaranteed to consist purely of \"field recordings\" as might be defined by practitioners. The intention is to represent the content of an archive collection on such a topic, rather than to represent a controlled definition of such a topic.\n\nEach audio file has a corresponding text file, containing metadata such as author and tags. \nThe dataset has been randomly split into 10 equal-size subsets. This is so that you can perform 10-fold crossvalidation in machine-learning experiments, or can use fixed subsets of the data (e.g. use one subset for development, and others for later validation). Each of the 10 subsets has about 128 minutes of audio; the dataset totals over 21 hours of audio.\n\n","metadata":{}},{"cell_type":"code","source":"# for no-call classification (using the freefield data)\n\n#credit to : \n\n\n#get all the json files (with description of the sounds)\nfile_list = Path(\"../input/freefield1010/freefield1010\").rglob(\"*.json\")\n\nall_audio = []\n\nfor filepath in file_list:\n    #conv json to pd \n    row = json_to_pd(filepath)\n    #add filepath to image()\n    row['filepath'] = str(filepath).rsplit('.',maxsplit=1)[0]  + '.wav'\n    \n    #append row to list\n    all_audio.append(row)\n    \n\n    \n    \nfreefield_df = pd.concat(all_audio,\n                         ignore_index=True)\n\n#check if there is a bird call in audio \nfreefield_df['has_bird_call'] = freefield_df['tags'].apply(lambda x: 'bird' in x).astype(int)\n\nfreefield_df.head(3)","metadata":{"execution":{"iopub.status.busy":"2022-05-02T08:23:11.431047Z","iopub.execute_input":"2022-05-02T08:23:11.431376Z","iopub.status.idle":"2022-05-02T08:24:11.777484Z","shell.execute_reply.started":"2022-05-02T08:23:11.431349Z","shell.execute_reply":"2022-05-02T08:24:11.776561Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#check number of birdcalls vs no calls in freefield data \nprint('Number of bird calls',freefield_df[freefield_df['has_bird_call']==1].shape[0])\nprint('Number of bird calls',freefield_df[freefield_df['has_bird_call']!=1].shape[0])\n\n","metadata":{"execution":{"iopub.status.busy":"2022-05-02T08:24:11.778633Z","iopub.execute_input":"2022-05-02T08:24:11.778867Z","iopub.status.idle":"2022-05-02T08:24:11.815520Z","shell.execute_reply.started":"2022-05-02T08:24:11.778841Z","shell.execute_reply":"2022-05-02T08:24:11.814466Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#undersampling the dataset, to have 1/4 of bird calls, 3/4 of no-bird-calls.\n\nfreefield_unsam = freefield_df[freefield_df['has_bird_call']==1].copy()\nfreefield_unsam = freefield_unsam.append(freefield_df[freefield_df['has_bird_call']!=1].sample(n= 378 * 3,random_state=RANDOM_SEED).copy(),\n                                        ignore_index=True)\n\nfreefield_unsam.head(2)","metadata":{"execution":{"iopub.status.busy":"2022-05-02T08:24:11.817053Z","iopub.execute_input":"2022-05-02T08:24:11.817461Z","iopub.status.idle":"2022-05-02T08:24:11.863579Z","shell.execute_reply.started":"2022-05-02T08:24:11.817418Z","shell.execute_reply":"2022-05-02T08:24:11.862992Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"freefield_unsam.shape","metadata":{"execution":{"iopub.status.busy":"2022-05-02T08:24:11.865339Z","iopub.execute_input":"2022-05-02T08:24:11.865667Z","iopub.status.idle":"2022-05-02T08:24:11.870836Z","shell.execute_reply.started":"2022-05-02T08:24:11.865639Z","shell.execute_reply":"2022-05-02T08:24:11.870010Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Extracting Spectrograms","metadata":{}},{"cell_type":"markdown","source":"**Helper Functions**","metadata":{}},{"cell_type":"code","source":"def plot_spec(path):\n    fig,ax = plt.subplots(figsize=(12,6))\n    \n    im = plt.imread(fname=path)\n    plt.axis('off')\n    plt.imshow(im,cmap='jet')\n    plt.colorbar(shrink=0.25)\n    plt.show()","metadata":{"execution":{"iopub.status.busy":"2022-05-02T08:50:25.148758Z","iopub.execute_input":"2022-05-02T08:50:25.149480Z","iopub.status.idle":"2022-05-02T08:50:25.156471Z","shell.execute_reply.started":"2022-05-02T08:50:25.149435Z","shell.execute_reply":"2022-05-02T08:50:25.155609Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Add Audio Augmentations**","metadata":{}},{"cell_type":"code","source":"#Audio Augmentation:\naugmentations = Compose(\n    [\n            FrequencyMask(min_frequency_band=0.05, max_frequency_band=0.15, p=0.25),\n            TimeStretch(min_rate=0.11,max_rate=0.3,p=0.25),\n            AddGaussianSNR(min_snr_in_db=5, max_snr_in_db=40, p=0.25)\n                        ]\n                        )","metadata":{"execution":{"iopub.status.busy":"2022-05-02T08:24:18.907253Z","iopub.execute_input":"2022-05-02T08:24:18.907777Z","iopub.status.idle":"2022-05-02T08:24:18.912659Z","shell.execute_reply.started":"2022-05-02T08:24:18.907743Z","shell.execute_reply":"2022-05-02T08:24:18.911887Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"\n#function to save stfts\n\ndef save_stft(file_path,\n              dir_path,\n              Augment=True,\n              ):\n    '''extracting mel-specs from given audio data and saving them to given folder'''\n    \n    #list to store spectogram and labels\n    stft_id = []\n        \n    #load audio     \n    sig,sr=librosa.load(file_path,\n                        sr=SAMPLE_RATE)\n    \n    \n    n=0\n    # break the signal into n second chunks\n    for i in range(0,len(sig),int(SIGNAL_LENGTH*SAMPLE_RATE)):\n        \n        window = sig[i:i + int(SIGNAL_LENGTH * SAMPLE_RATE)]\n\n        # End of signal\n        if len(window) < int(SIGNAL_LENGTH * SAMPLE_RATE):\n            break\n            \n            \n            \n        #Apply audio Augmentations :   \n        if Augment:\n            window=augmentations(window,\n                                 sample_rate=sr)\n            \n        # extracting mel-spectrograms:\n        mel_spec = librosa.feature.melspectrogram(window,\n                                                  sr=SAMPLE_RATE, \n                                                  n_fft=1024, \n                                                  hop_length=hop_length, \n                                                  n_mels=SPEC_SHAPE[0], \n                                                  fmin=FMIN, \n                                                  fmax=FMAX\n                                                  )\n        \n        \n        # log scaling (convert to decibels)\n        mel_spec = librosa.core.power_to_db(mel_spec,\n                                            ref=np.max)\n        \n        # Normalize\n        mel_spec -= mel_spec.min()\n        mel_spec /= mel_spec.max()\n        \n        \n        #saving Image\n        \n        #image_id\n        ids=file_path.split('/')[-1].split('.')[0]\n        \n        save_id=f'{ids}_{n}.jpg'\n        save_path=os.path.join(dir_path,save_id)\n        \n        n+=1\n        \n        image = Image.fromarray(mel_spec * 255.0).convert(\"L\")\n        image.save(save_path)\n        \n        #saving_image ids and labels\n        stft_id.append(save_id)\n        \n    return stft_id\n\n","metadata":{"execution":{"iopub.status.busy":"2022-05-02T08:24:19.328770Z","iopub.execute_input":"2022-05-02T08:24:19.329268Z","iopub.status.idle":"2022-05-02T08:24:19.341819Z","shell.execute_reply.started":"2022-05-02T08:24:19.329231Z","shell.execute_reply":"2022-05-02T08:24:19.340665Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# **Extracting Train Spectrograms**","metadata":{}},{"cell_type":"code","source":"train_ids=[]\nfor idx,row in tqdm(train_metadata.iterrows()):\n    \n    #save spectograms\n    audio_ids = save_stft(file_path= row.filepath,\n                           dir_path =Train_Spectrograms)\n    \n    train_ids.extend(audio_ids)\n    ","metadata":{"execution":{"iopub.status.busy":"2022-05-02T08:31:56.427448Z","iopub.execute_input":"2022-05-02T08:31:56.428095Z","iopub.status.idle":"2022-05-02T08:32:07.429159Z","shell.execute_reply.started":"2022-05-02T08:31:56.428059Z","shell.execute_reply":"2022-05-02T08:32:07.428329Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Saving DataFRame with spectrogram ids**","metadata":{}},{"cell_type":"code","source":"train_df = pd.DataFrame(train_ids,\n                        columns=['spec_id'])\ntrain_df['file_id'] = train_df['spec_id'].apply(lambda x: x.split('_')[0])\n\n\ntrain_metadata['file_id'] = train_metadata['filename'].apply(lambda x: x.split('/')[1].split('.')[0])\n\n\n#join dfs\ntrain_df = train_df.merge(train_metadata,\n                          on='file_id',\n                          how='left')\n\n\ntrain_df.to_csv('train_df.csv',\n                index=False)\n\ntrain_df.shape","metadata":{"execution":{"iopub.status.busy":"2022-05-02T08:43:58.629278Z","iopub.execute_input":"2022-05-02T08:43:58.629618Z","iopub.status.idle":"2022-05-02T08:43:58.668976Z","shell.execute_reply.started":"2022-05-02T08:43:58.629586Z","shell.execute_reply":"2022-05-02T08:43:58.668391Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plot_spec('./Train_Spectrograms/XC177993_0.jpg')","metadata":{"execution":{"iopub.status.busy":"2022-05-02T08:51:46.187846Z","iopub.execute_input":"2022-05-02T08:51:46.188183Z","iopub.status.idle":"2022-05-02T08:51:46.402783Z","shell.execute_reply.started":"2022-05-02T08:51:46.188150Z","shell.execute_reply":"2022-05-02T08:51:46.401879Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plot_spec('./Train_Spectrograms/XC177993_1.jpg')","metadata":{"execution":{"iopub.status.busy":"2022-05-02T08:51:20.506922Z","iopub.execute_input":"2022-05-02T08:51:20.507681Z","iopub.status.idle":"2022-05-02T08:51:20.719964Z","shell.execute_reply.started":"2022-05-02T08:51:20.507633Z","shell.execute_reply":"2022-05-02T08:51:20.719103Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# **Extracting Spectrograms from free-field data.(for no-call classifier)**","metadata":{}},{"cell_type":"code","source":"freefield_ids=[]\nfor idx, row in tqdm(freefield_unsam.iterrows()):\n    \n    #save spectograms\n    audio_ids = save_stft(file_path= row.filepath,\n                           dir_path =Freefield_Spectograms)\n    \n    freefield_ids.extend(audio_ids)\n    \nprint('Number of samples extracted ',len(freefield_ids))","metadata":{"execution":{"iopub.status.busy":"2022-05-02T08:40:59.477555Z","iopub.execute_input":"2022-05-02T08:40:59.478388Z","iopub.status.idle":"2022-05-02T08:41:06.050830Z","shell.execute_reply.started":"2022-05-02T08:40:59.478344Z","shell.execute_reply":"2022-05-02T08:41:06.042438Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plot_spec('./Freefield_Spectrograms/44218_0.jpg')","metadata":{"execution":{"iopub.status.busy":"2022-05-02T08:52:24.907792Z","iopub.execute_input":"2022-05-02T08:52:24.908455Z","iopub.status.idle":"2022-05-02T08:52:25.129026Z","shell.execute_reply.started":"2022-05-02T08:52:24.908417Z","shell.execute_reply":"2022-05-02T08:52:25.128118Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plot_spec('./Freefield_Spectrograms/111096_0.jpg')","metadata":{"execution":{"iopub.status.busy":"2022-05-02T08:53:16.227626Z","iopub.execute_input":"2022-05-02T08:53:16.228213Z","iopub.status.idle":"2022-05-02T08:53:16.443455Z","shell.execute_reply.started":"2022-05-02T08:53:16.228176Z","shell.execute_reply":"2022-05-02T08:53:16.442546Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#saving free field df \nfreefield_unsam.to_csv('freefield_downsampled.csv',index=False)","metadata":{"execution":{"iopub.status.busy":"2022-05-02T08:41:18.517793Z","iopub.execute_input":"2022-05-02T08:41:18.518112Z","iopub.status.idle":"2022-05-02T08:41:18.597022Z","shell.execute_reply.started":"2022-05-02T08:41:18.518078Z","shell.execute_reply":"2022-05-02T08:41:18.596159Z"},"trusted":true},"execution_count":null,"outputs":[]}]}