{"cells":[{"metadata":{},"cell_type":"markdown","source":"# Exploring a database of wildlife recordings\n\nThis database contains a series of 1 minute recordings containing at least one call of known a wildlife species. These are identified in the file `train_tp.csv` with start and end times and frequency bands. \n\nEach recording may also contain calls of other species. Some of them are marked in the file `train_fp.csv` as being wrongly identified by some automated algorithm.\n\nHere we explore the times and frequency bands of each species and pinpoint them in a representation of the recording. "},{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true},"cell_type":"code","source":"\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n\n","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"We'll use librosa to read audio and perform some analysis"},{"metadata":{"_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","trusted":true},"cell_type":"code","source":"import librosa as lr\nimport librosa.display as lrd\nimport os\nimport matplotlib.pyplot as plt\nimport matplotlib.patches as patches\nimport seaborn as sns","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Let's read both tables to identify true and false positives on the spectrum"},{"metadata":{"trusted":true},"cell_type":"code","source":"tpdf = pd.read_csv('/kaggle/input/rfcx-species-audio-detection/train_tp.csv')\nfpdf = pd.read_csv('/kaggle/input/rfcx-species-audio-detection/train_fp.csv')\n","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Number of samples per call type and species"},{"metadata":{"trusted":true},"cell_type":"code","source":"tpdf['duration'] = tpdf.t_max-tpdf.t_min\ntpdf['bandwidth'] = tpdf.f_max-tpdf.f_min\n\ntpdf.pivot_table(index='species_id',columns='songtype_id',\n                 values='duration',aggfunc='count').fillna(0)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Metadata per species\n\nNotice how the duration and frequency bands are highly characteristic of each species. \n\nThe trouble will be to establish the event boundaries in new sound files"},{"metadata":{},"cell_type":"markdown","source":"### Call duration\n"},{"metadata":{"trusted":true},"cell_type":"code","source":"fig,ax = plt.subplots(1,figsize=(10,4))\nsns.boxplot(data=tpdf,y='duration',x='species_id',hue='songtype_id')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### Lower frequency of call band"},{"metadata":{"trusted":true},"cell_type":"code","source":"fig,ax = plt.subplots(1,figsize=(10,4))\nsns.boxplot(data=tpdf,y='f_min',x='species_id',hue='songtype_id')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### Uper frequency of call band"},{"metadata":{"trusted":true},"cell_type":"code","source":"fig,ax = plt.subplots(1,figsize=(10,4))\nsns.boxplot(data=tpdf,y='f_max',x='species_id',hue='songtype_id')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### Bandwidth"},{"metadata":{"trusted":true},"cell_type":"code","source":"fig,ax = plt.subplots(1,figsize=(10,4))\nsns.boxplot(data=tpdf,y='bandwidth',x='species_id',hue='songtype_id')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Sample audio\n\nNow let's select a sample audio file and mark the events \n\nWe'll use the MEL spectrum to visualize the audio data. It's a sort of spectrogram where the frequency bands are logarithmic"},{"metadata":{"trusted":true},"cell_type":"code","source":"row=tpdf.sample()\n\nbase_dir = '/kaggle/input/rfcx-species-audio-detection/train/'\nw,sr = lr.load(os.path.join(base_dir,row.iloc[0]['recording_id']+'.flac'))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"fmax_mel=11000\nhop_length=256\nn_wind=1024\nms = lr.feature.melspectrogram(w,sr,n_fft=n_wind,hop_length=hop_length,fmax=fmax_mel)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"fig,ax= plt.subplots(1)\n\n# convert to dB because the range is too wide\ndbs = 20*np.log10(ms)\n\n\n# This shows the MEL spectrogram\nimg = lrd.specshow(dbs,sr=sr,fmax=fmax_mel,\n                   y_axis=\"mel\",x_axis=\"time\",\n                   hop_length=hop_length,\n                   cmap='gray')\n\n# Plot the True positive events on this file\nrecrows = tpdf[tpdf.recording_id==row.iloc[0].recording_id]\nprint(f'Number of TRUE positives: {len(recrows)}')\nfor ir,rrow in recrows.iterrows():\n    rect = patches.Rectangle((rrow.t_min,rrow.f_min),\n                             rrow.t_max-rrow.t_min,rrow.f_max-rrow.f_min,\n                             linewidth=1,edgecolor='g',facecolor='g',alpha=.2)\n    #plt.axvspan(rrow.t_min,rrow.t_max,color='g',alpha=.2)\n    ax.add_patch(rect)\n    \n# Plot the False Positives on this file\nrecrows = fpdf[fpdf.recording_id==row.iloc[0].recording_id]\nprint(f'Number of FALSE positives: {len(recrows)}')\nfor ir,rrow in recrows.iterrows():\n    rect = patches.Rectangle((rrow.t_min,rrow.f_min),\n                             rrow.t_max-rrow.t_min,rrow.f_max-rrow.f_min,\n                             linewidth=1,edgecolor='r',facecolor='r',alpha=.2)\n    #plt.axvspan(rrow.t_min,rrow.t_max,color='g',alpha=.2)\n    ax.add_patch(rect)\n","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Notice how some of the true events are hard to distinguish from the background noise"},{"metadata":{},"cell_type":"markdown","source":"## Extract a normalised spectrum"},{"metadata":{"trusted":true},"cell_type":"code","source":"avs = np.median(dbs,axis=1)\ndvs = np.diff(np.percentile(dbs,[25,75],axis=1),axis=0)[0]\n\nzs = (dbs-np.tile(avs[:,np.newaxis],(1,dbs.shape[1])))/np.tile(dvs[:,np.newaxis],(1,dbs.shape[1]))\nimg = lrd.specshow(zs,sr=sr,fmax=fmax_mel,\n                   y_axis=\"mel\",x_axis=\"time\",\n                   hop_length=hop_length,\n                   cmap='gray')\n","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"zzs = zs.copy()\nzzs[zzs<1] = 0\n\nfig,ax= plt.subplots(1,figsize=(12,4))\n\n\nimg = lrd.specshow(zzs,sr=sr,fmax=fmax_mel,\n                   y_axis=\"mel\",x_axis=\"time\",\n                   hop_length=hop_length,\n                   cmap='gray')\n\n# Plot the True positive events on this file\ntrecrows = tpdf[tpdf.recording_id==row.iloc[0].recording_id]\nprint(f'Number of TRUE positives: {len(trecrows)}')\nfor ir,rrow in trecrows.iterrows():\n    rect = patches.Rectangle((rrow.t_min,rrow.f_min),\n                             rrow.t_max-rrow.t_min,rrow.f_max-rrow.f_min,\n                             linewidth=1,edgecolor='g',facecolor='g',alpha=.5)\n    #plt.axvspan(rrow.t_min,rrow.t_max,color='g',alpha=.2)\n    ax.add_patch(rect)\n    \n# Plot the False Positives on this file\nfrecrows = fpdf[fpdf.recording_id==row.iloc[0].recording_id]\nprint(f'Number of FALSE positives: {len(frecrows)}')\nfor ir,rrow in frecrows.iterrows():\n    rect = patches.Rectangle((rrow.t_min,rrow.f_min),\n                             rrow.t_max-rrow.t_min,rrow.f_max-rrow.f_min,\n                             linewidth=1,edgecolor='r',facecolor='r',alpha=.2)\n    #plt.axvspan(rrow.t_min,rrow.t_max,color='g',alpha=.5)\n    ax.add_patch(rect)\n","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"t=np.linspace(0,len(w-n_wind)/sr,dbs.shape[1])\n\nfor ir,rrow in trecrows.iterrows():\n    idx = (t>rrow.t_min) & (t<rrow.t_max)\n    fig,ax = plt.subplots(1)\n    img = lrd.specshow(zzs[:,idx],sr=sr,fmax=fmax_mel,\n                   y_axis=\"mel\",x_axis=\"time\",\n                   hop_length=hop_length,\n                   cmap='gray')\n    ax.axhspan(rrow.f_min,rrow.f_max,color='g',alpha=.4)\nfor ir,rrow in frecrows.iterrows():\n    idx = (t>rrow.t_min) & (t<rrow.t_max)\n    fig,ax = plt.subplots(1)\n    img = lrd.specshow(zzs[:,idx],sr=sr,fmax=fmax_mel,\n                   y_axis=\"mel\",x_axis=\"time\",\n                   hop_length=hop_length,\n                   cmap='gray')\n    ax.axhspan(rrow.f_min,rrow.f_max,color='r',alpha=.4)\n","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Compare a few samples of the same species"},{"metadata":{"trusted":true},"cell_type":"code","source":"srows = tpdf[(tpdf.species_id==row.iloc[0].species_id) & (tpdf.songtype_id==row.iloc[0].songtype_id)]\nsrows.sample(10)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"def get_ev_zspec(row):\n    w,sr = lr.load(os.path.join(base_dir,row['recording_id']+'.flac'))\n    ms = lr.feature.melspectrogram(w,sr,n_fft=n_wind,hop_length=hop_length,fmax=fmax_mel)    \n    dbs = 20*np.log10(ms)\n    avs = np.median(dbs,axis=1)\n    dvs = np.diff(np.percentile(dbs,[25,75],axis=1),axis=0)[0]\n\n    zs = (dbs-np.tile(avs[:,np.newaxis],(1,dbs.shape[1])))\n    zd = np.tile(dvs[:,np.newaxis],(1,dbs.shape[1]))\n    t=np.linspace(0,len(w-n_wind)/sr,dbs.shape[1])\n\n    idx = (t>row.t_min) & (t<row.t_max)\n    zs[zs<zd]=0\n    return zs[:,idx]\n","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"nsam = 6\nncols = 3\nnrows = int(np.ceil(nsam/ncols))\n\nfig,ax = plt.subplots(nrows,ncols,sharex=True,sharey=True)\naxf = ax.flatten()\n\nfor ii, (ir, srw) in enumerate(srows.sample(nsam).iterrows()):\n    zss = get_ev_zspec(srw)\n    img = lrd.specshow(zss,sr=sr,fmax=fmax_mel,\n               y_axis=\"mel\",x_axis=\"time\",\n               hop_length=hop_length,\n               cmap='gray',ax=axf[ii])\n    axf[ii].axhspan(srw.f_min,srw.f_max,color='g',alpha=.4)\n\n    ","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat":4,"nbformat_minor":4}