{
  "id": 197750,
  "title": "How to deal with audio detection ?",
  "url": "/competitions/rfcx-species-audio-detection/discussion/197750",
  "author_name": "",
  "post_date": "2020-11-17T22:11:21.099520400Z",
  "votes": 21,
  "comment_count": 5,
  "views": 0,
  "content": "<p>It was a great pleasure for me to compete in <a href=\"https://www.kaggle.com/c/birdsong-recognition\" target=\"_blank\">the latest Kaggle audio detection  competition</a> which is not too different from this one: definitely, audio detection is now drawing more attention on Kaggle than ever.</p>\n<p>Based on my last experience, I will be sharing some common techniques that could be used to deal with audio data classification/labeling.</p>\n<h2>Recurrent NN models</h2>\n<p>The first idea could be to treat the audio records as simple time series. Hence, any recurrent NN could be used (LSTM, GRU, TCN, …). That could be a great way even if many SOTA audio detection/classification models don't go that way</p>\n<h2>From audio to image :  audio detection as computer vision problem</h2>\n<p>The most used approach when dealing with audio data is this one. The waveforms are transformed into images using, for example MFCCs. Those images are fed to a computer vision model (EfficientNet, ResNet, ResNext …). Hence, audio  detection can fully leverage all the advancements in computer vision.</p>\n<p>Hopefully, I will be sharing some kernels here whenever I got some time off.</p>\n<h2># Update 1</h2>\n<p>As promised, I've <a href=\"https://www.kaggle.com/kneroma/inference-resnest-rfcx-audio-detection\" target=\"_blank\">released my first model</a>. It's clearly a more interesting baseline for next RFCX audio detection models (I'm sure there will be plenty of models since the dataset seems very clean and not too big)</p>\n<blockquote>\n  <p>To be continued</p>\n</blockquote>",
  "messages": [
    {
      "id": "1082432",
      "postDate": "11/17/2020 22:11:21",
      "content": "<p>It was a great pleasure for me to compete in <a href=\"https://www.kaggle.com/c/birdsong-recognition\" target=\"_blank\">the latest Kaggle audio detection  competition</a> which is not too different from this one: definitely, audio detection is now drawing more attention on Kaggle than ever.</p>\n<p>Based on my last experience, I will be sharing some common techniques that could be used to deal with audio data classification/labeling.</p>\n<h2>Recurrent NN models</h2>\n<p>The first idea could be to treat the audio records as simple time series. Hence, any recurrent NN could be used (LSTM, GRU, TCN, …). That could be a great way even if many SOTA audio detection/classification models don't go that way</p>\n<h2>From audio to image :  audio detection as computer vision problem</h2>\n<p>The most used approach when dealing with audio data is this one. The waveforms are transformed into images using, for example MFCCs. Those images are fed to a computer vision model (EfficientNet, ResNet, ResNext …). Hence, audio  detection can fully leverage all the advancements in computer vision.</p>\n<p>Hopefully, I will be sharing some kernels here whenever I got some time off.</p>\n<h2># Update 1</h2>\n<p>As promised, I've <a href=\"https://www.kaggle.com/kneroma/inference-resnest-rfcx-audio-detection\" target=\"_blank\">released my first model</a>. It's clearly a more interesting baseline for next RFCX audio detection models (I'm sure there will be plenty of models since the dataset seems very clean and not too big)</p>\n<blockquote>\n  <p>To be continued</p>\n</blockquote>",
      "rawMarkdown": "It was a great pleasure for me to compete in [the latest Kaggle audio detection  competition](https://www.kaggle.com/c/birdsong-recognition) which is not too different from this one: definitely, audio detection is now drawing more attention on Kaggle than ever.\n\n Based on my last experience, I will be sharing some common techniques that could be used to deal with audio data classification/labeling.\n\n## Recurrent NN models\nThe first idea could be to treat the audio records as simple time series. Hence, any recurrent NN could be used (LSTM, GRU, TCN, ...). That could be a great way even if many SOTA audio detection/classification models don't go that way\n\n## From audio to image :  audio detection as computer vision problem\nThe most used approach when dealing with audio data is this one. The waveforms are transformed into images using, for example MFCCs. Those images are fed to a computer vision model (EfficientNet, ResNet, ResNext ...). Hence, audio  detection can fully leverage all the advancements in computer vision.\n\nHopefully, I will be sharing some kernels here whenever I got some time off.\n\n\n## # Update 1\nAs promised, I've [released my first model](https://www.kaggle.com/kneroma/inference-resnest-rfcx-audio-detection). It's clearly a more interesting baseline for next RFCX audio detection models (I'm sure there will be plenty of models since the dataset seems very clean and not too big)\n\n\n> To be continued",
      "votes": null
    },
    {
      "id": "1091288",
      "postDate": "11/25/2020 23:10:09",
      "content": "<p>thanks for your sharing! I have a basic question about how the problem is framed and its target. I saw in your kernel that seems like you were training on a random crop of an audio clip, but in <code>train_tp.csv</code> it shows the target is only present in a period of time i.e. <code>t_min</code> and <code>t_max</code>, so wouldn't a random cropping likely miss this period of clip and thus the target should be null or something? Couldn't we just feed the entire audio clip to cnn as input instead of small random crops?   </p>",
      "rawMarkdown": "thanks for your sharing! I have a basic question about how the problem is framed and its target. I saw in your kernel that seems like you were training on a random crop of an audio clip, but in `train_tp.csv` it shows the target is only present in a period of time i.e. `t_min` and `t_max`, so wouldn't a random cropping likely miss this period of clip and thus the target should be null or something? Couldn't we just feed the entire audio clip to cnn as input instead of small random crops?",
      "votes": null
    },
    {
      "id": "1091313",
      "postDate": "11/25/2020 23:43:49",
      "content": "<p>We could feed the entire audio for sure. It's up to you :) .<br>\nBut, if you're cropping, its means that the target will depend on the cropped chunk. For example, if you crop  a chunk which overlap with zero (t_min, t_max), all 24 targets must be set to zeros.</p>",
      "rawMarkdown": "We could feed the entire audio for sure. It's up to you :) .\nBut, if you're cropping, its means that the target will depend on the cropped chunk. For example, if you crop  a chunk which overlap with zero (t_min, t_max), all 24 targets must be set to zeros.",
      "votes": null
    },
    {
      "id": "1091373",
      "postDate": "11/26/2020 01:01:39",
      "content": "<p>I see… Just want to confirm, in your shared notebook I'm a bit confused that you set the specie_id to empty list: <code>data[\"species_id\"] = [[] for _ in range(len(data))]</code> but in your dataset you get the target by setting the index of the class (the specie id) to 1, is this because the kernel is for inference, and for training we should match the cropped chunk to the target period accordingly based on <code>train_tp</code>? </p>",
      "rawMarkdown": "I see... Just want to confirm, in your shared notebook I'm a bit confused that you set the specie_id to empty list: `data[\"species_id\"] = [[] for _ in range(len(data))]` but in your dataset you get the target by setting the index of the class (the specie id) to 1, is this because the kernel is for inference, and for training we should match the cropped chunk to the target period accordingly based on `train_tp`?",
      "votes": null
    },
    {
      "id": "1091410",
      "postDate": "11/26/2020 02:07:50",
      "content": "<p>Yes, It's because It's inference. I need not the species IDs (we're supposed to predict them for the test set), I just put them there to make my code work.</p>",
      "rawMarkdown": "Yes, It's because It's inference. I need not the species IDs (we're supposed to predict them for the test set), I just put them there to make my code work.",
      "votes": null
    },
    {
      "id": "1105535",
      "postDate": "12/08/2020 01:27:46",
      "content": "<p><a href=\"https://www.kaggle.com/kneroma\" target=\"_blank\">@kneroma</a> thanks for sharing! I got a question similar to <a href=\"https://www.kaggle.com/samshipengs\" target=\"_blank\">@samshipengs</a>. Let's say we take advantage of <code>t_min</code> and <code>t_max</code> to better detect species rather than using random crops, for example, by using small time windows of 10 seconds which start at <code>t_min</code> and contain <code>t_max</code>. </p>\n<p>Then when it comes to the test set which has audios of 60 seconds with no <code>t_min</code> nor <code>_t_max</code> information, should we randomly take 10 second clips and test on those? Should we take 6 audio samples of 10 seconds each and average predictions for each sample? How do you tackle this issue?</p>\n<p>I posted a similar question <a href=\"https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/201827\" target=\"_blank\">here</a></p>\n<p>Thanks in advance!</p>",
      "rawMarkdown": "kneroma thanks for sharing! I got a question similar to @samshipengs. Let's say we take advantage of `t_min` and `t_max` to better detect species rather than using random crops, for example, by using small time windows of 10 seconds which start at `t_min` and contain `t_max`. \n\nThen when it comes to the test set which has audios of 60 seconds with no `t_min` nor `_t_max` information, should we randomly take 10 second clips and test on those? Should we take 6 audio samples of 10 seconds each and average predictions for each sample? How do you tackle this issue?\n\nI posted a similar question [here](https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/201827)\n\nThanks in advance!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1091288,
      "author_name": "samshipengs",
      "author_url": "",
      "post_date": "11/25/2020 23:10:09",
      "content": "<p>thanks for your sharing! I have a basic question about how the problem is framed and its target. I saw in your kernel that seems like you were training on a random crop of an audio clip, but in <code>train_tp.csv</code> it shows the target is only present in a period of time i.e. <code>t_min</code> and <code>t_max</code>, so wouldn't a random cropping likely miss this period of clip and thus the target should be null or something? Couldn't we just feed the entire audio clip to cnn as input instead of small random crops?   </p>",
      "votes": null,
      "replies": [
        {
          "id": 1091313,
          "author_name": "kneroma",
          "author_url": "",
          "post_date": "11/25/2020 23:43:49",
          "content": "<p>We could feed the entire audio for sure. It's up to you :) .<br>\nBut, if you're cropping, its means that the target will depend on the cropped chunk. For example, if you crop  a chunk which overlap with zero (t_min, t_max), all 24 targets must be set to zeros.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1091373,
          "author_name": "samshipengs",
          "author_url": "",
          "post_date": "11/26/2020 01:01:39",
          "content": "<p>I see… Just want to confirm, in your shared notebook I'm a bit confused that you set the specie_id to empty list: <code>data[\"species_id\"] = [[] for _ in range(len(data))]</code> but in your dataset you get the target by setting the index of the class (the specie id) to 1, is this because the kernel is for inference, and for training we should match the cropped chunk to the target period accordingly based on <code>train_tp</code>? </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1091410,
          "author_name": "kneroma",
          "author_url": "",
          "post_date": "11/26/2020 02:07:50",
          "content": "<p>Yes, It's because It's inference. I need not the species IDs (we're supposed to predict them for the test set), I just put them there to make my code work.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1105535,
      "author_name": "alejopaullier",
      "author_url": "",
      "post_date": "12/08/2020 01:27:46",
      "content": "<p><a href=\"https://www.kaggle.com/kneroma\" target=\"_blank\">@kneroma</a> thanks for sharing! I got a question similar to <a href=\"https://www.kaggle.com/samshipengs\" target=\"_blank\">@samshipengs</a>. Let's say we take advantage of <code>t_min</code> and <code>t_max</code> to better detect species rather than using random crops, for example, by using small time windows of 10 seconds which start at <code>t_min</code> and contain <code>t_max</code>. </p>\n<p>Then when it comes to the test set which has audios of 60 seconds with no <code>t_min</code> nor <code>_t_max</code> information, should we randomly take 10 second clips and test on those? Should we take 6 audio samples of 10 seconds each and average predictions for each sample? How do you tackle this issue?</p>\n<p>I posted a similar question <a href=\"https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/201827\" target=\"_blank\">here</a></p>\n<p>Thanks in advance!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1082432": "It was a great pleasure for me to compete in [the latest Kaggle audio detection  competition](https://www.kaggle.com/c/birdsong-recognition) which is not too different from this one: definitely, audio detection is now drawing more attention on Kaggle than ever.\n\n Based on my last experience, I will be sharing some common techniques that could be used to deal with audio data classification/labeling.\n\n## Recurrent NN models\nThe first idea could be to treat the audio records as simple time series. Hence, any recurrent NN could be used (LSTM, GRU, TCN, ...). That could be a great way even if many SOTA audio detection/classification models don't go that way\n\n## From audio to image :  audio detection as computer vision problem\nThe most used approach when dealing with audio data is this one. The waveforms are transformed into images using, for example MFCCs. Those images are fed to a computer vision model (EfficientNet, ResNet, ResNext ...). Hence, audio  detection can fully leverage all the advancements in computer vision.\n\nHopefully, I will be sharing some kernels here whenever I got some time off.\n\n\n## # Update 1\nAs promised, I've [released my first model](https://www.kaggle.com/kneroma/inference-resnest-rfcx-audio-detection). It's clearly a more interesting baseline for next RFCX audio detection models (I'm sure there will be plenty of models since the dataset seems very clean and not too big)\n\n\n> To be continued",
    "1091288": "thanks for your sharing! I have a basic question about how the problem is framed and its target. I saw in your kernel that seems like you were training on a random crop of an audio clip, but in `train_tp.csv` it shows the target is only present in a period of time i.e. `t_min` and `t_max`, so wouldn't a random cropping likely miss this period of clip and thus the target should be null or something? Couldn't we just feed the entire audio clip to cnn as input instead of small random crops?",
    "1091313": "We could feed the entire audio for sure. It's up to you :) .\nBut, if you're cropping, its means that the target will depend on the cropped chunk. For example, if you crop  a chunk which overlap with zero (t_min, t_max), all 24 targets must be set to zeros.",
    "1091373": "I see... Just want to confirm, in your shared notebook I'm a bit confused that you set the specie_id to empty list: `data[\"species_id\"] = [[] for _ in range(len(data))]` but in your dataset you get the target by setting the index of the class (the specie id) to 1, is this because the kernel is for inference, and for training we should match the cropped chunk to the target period accordingly based on `train_tp`?",
    "1091410": "Yes, It's because It's inference. I need not the species IDs (we're supposed to predict them for the test set), I just put them there to make my code work.",
    "1105535": "kneroma thanks for sharing! I got a question similar to @samshipengs. Let's say we take advantage of `t_min` and `t_max` to better detect species rather than using random crops, for example, by using small time windows of 10 seconds which start at `t_min` and contain `t_max`. \n\nThen when it comes to the test set which has audios of 60 seconds with no `t_min` nor `_t_max` information, should we randomly take 10 second clips and test on those? Should we take 6 audio samples of 10 seconds each and average predictions for each sample? How do you tackle this issue?\n\nI posted a similar question [here](https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/201827)\n\nThanks in advance!"
  },
  "source": "meta"
}