{
  "id": 217460,
  "title": "Yet another way to crop audio with Tensorflow",
  "url": "/competitions/rfcx-species-audio-detection/discussion/217460",
  "author_name": "",
  "post_date": "2021-02-06T21:01:51.241894300Z",
  "votes": 14,
  "comment_count": 4,
  "views": 0,
  "content": "<p>When I started looking at this competition one thing that confused me a little was how to effectively crop the audio, especially using Tensorflow to maximize the efficiency, I studied a few public notebooks and ended up making a lot of experiments into that, and have written a few functions that should be able to crop the data with some flexibility.<br>\nBesides that, the importance of audio cropping has been discussed many times here, especially for creating an effective validation, references: <a href=\"https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/200922\" target=\"_blank\">[1]</a>, <a href=\"https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/207624\" target=\"_blank\">[2]</a>, <a href=\"https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/215233\" target=\"_blank\">[3]</a>, <a href=\"https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/216564\" target=\"_blank\">[4]</a>.</p>\n<p>The Idea is to first crop the audio with a given <code>crop size</code> then apply <strong>random crops</strong> at this <strong>first crop</strong>. The idea is very similar to what most public notebooks uses, one thing that I tried to make sure of is that the first crop had <strong>reasonable padding</strong> in cases where the label was small.</p>\n<p>You can find the notebook here: <a href=\"https://www.kaggle.com/dimitreoliveira/rainforest-audio-classification-tf-improved\" target=\"_blank\">Rainforest-Audio classification TF Improved</a></p>\n<p>To help with the understanding let's look at these examples:</p>\n<p><strong>1st example</strong><br>\nThe first crop can have equal paddings on both sides.<br>\n<img src=\"https://raw.githubusercontent.com/dimitreOliveira/MachineLearning/master/Kaggle/Rainforest%20Connection%20Species%20Audio%20Detection/Audio%20crop%20diagram_2.png\" alt=\"\"></p>\n<p>The first crop can't have equal paddings, so one side will have larger padding.<br>\n<img src=\"https://raw.githubusercontent.com/dimitreOliveira/MachineLearning/master/Kaggle/Rainforest%20Connection%20Species%20Audio%20Detection/Audio%20crop%20diagram.png\" alt=\"\"></p>\n<hr>\n<h3>The code</h3>\n<pre><code>def crop_audio(audio, tmin, tmax, min_crop_size=MIN_CROP_SIZE, sample_rate=48000, max_size=60.):\n    \"\"\"\n        Crops a 'waveform' file to have {min_crop_size} size given, {tmin}, {tmax}, {sample_rate} and {max_size}.\n    \"\"\"\n    label_size = tmax - tmin\n\n    if label_size &gt;= min_crop_size: # No padding needed\n        cut_min = tmin\n        cut_max = tmax        \n    else: # Needs padding\n        pad_start = (min_crop_size - label_size) / 2\n        cut_min = tf.maximum(0., (tmin - pad_start))\n\n        pad_end = (min_crop_size - (label_size - cut_min))\n        cut_max = tf.minimum(max_size, (tmax + pad_end))\n\n        cut_size = cut_max - cut_min\n\n        if cut_size &lt; min_crop_size:\n            cut_min = tf.maximum(0., (cut_min - (min_crop_size - cut_size)))\n\n    cut_size = cut_max - cut_min\n\n    # Casting tensors\n    cut_min = tf.cast((cut_min * sample_rate), tf.int32)\n    cut_max = tf.cast((cut_max * sample_rate), tf.int32)\n    cut_size = tf.cast((cut_size * sample_rate), tf.int32)\n\n    audio = audio[cut_min:cut_max] # croping the audio\n    audio = audio[:cut_size] # making sure it has the max size\n\n    audio = tf.reshape(audio, [cut_size]) # making sure it has the expected shape\n    return audio\n</code></pre>\n<pre><code>def random_crop_audio(audio, crop_size=CROP_SIZE, sample_rate=48000, max_size=60):\n    \"\"\"\n        Randomly crops a 'waveform' file to have {crop_size} size given, {sample_rate} and {max_size}.\n    \"\"\"\n    start = tf.random.uniform([], minval=0, \n                              maxval=(max_size - crop_size), \n                              dtype=tf.int32)\n    cut_min = start * sample_rate\n    cut_max = (start + crop_size) * sample_rate\n\n    audio_size = len(audio)\n    if cut_max &gt; audio_size:\n        cut_min -= cut_max - audio_size\n        cut_max = cut_min + (crop_size * sample_rate)\n\n    # Casting tensors\n    cut_min = tf.cast(cut_min, tf.int32)\n    cut_max = tf.cast(cut_max, tf.int32)\n    cut_size = tf.cast((crop_size*sample_rate), tf.int32)\n\n    audio = audio[cut_min:cut_max] # croping the audio\n    audio = audio[:cut_size] # making sure it has the max size\n\n    audio = tf.reshape(audio, [cut_size]) # making sure it has the expected shape\n    return audio\n</code></pre>\n<hr>\n<p>I am not very familiar with audio data but I will add a few observations here</p>\n<ul>\n<li>For me <code>EfficientNet</code> seems to work better than <code>ResNets</code>, I could not get ResNets to perform as well as public notebooks.</li>\n<li>Mel-spectrograms performed worse than regular spectrograms, but I might be doing something wrong.</li>\n<li>Tweaking the crop sizes seems more relevant than the spectrograms resolution (height x width)</li>\n<li>Incrementing the architecture at the top of the backbone matters very little.</li>\n</ul>",
  "messages": [
    {
      "id": "1189261",
      "postDate": "02/06/2021 21:01:51",
      "content": "<p>When I started looking at this competition one thing that confused me a little was how to effectively crop the audio, especially using Tensorflow to maximize the efficiency, I studied a few public notebooks and ended up making a lot of experiments into that, and have written a few functions that should be able to crop the data with some flexibility.<br>\nBesides that, the importance of audio cropping has been discussed many times here, especially for creating an effective validation, references: <a href=\"https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/200922\" target=\"_blank\">[1]</a>, <a href=\"https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/207624\" target=\"_blank\">[2]</a>, <a href=\"https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/215233\" target=\"_blank\">[3]</a>, <a href=\"https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/216564\" target=\"_blank\">[4]</a>.</p>\n<p>The Idea is to first crop the audio with a given <code>crop size</code> then apply <strong>random crops</strong> at this <strong>first crop</strong>. The idea is very similar to what most public notebooks uses, one thing that I tried to make sure of is that the first crop had <strong>reasonable padding</strong> in cases where the label was small.</p>\n<p>You can find the notebook here: <a href=\"https://www.kaggle.com/dimitreoliveira/rainforest-audio-classification-tf-improved\" target=\"_blank\">Rainforest-Audio classification TF Improved</a></p>\n<p>To help with the understanding let's look at these examples:</p>\n<p><strong>1st example</strong><br>\nThe first crop can have equal paddings on both sides.<br>\n<img src=\"https://raw.githubusercontent.com/dimitreOliveira/MachineLearning/master/Kaggle/Rainforest%20Connection%20Species%20Audio%20Detection/Audio%20crop%20diagram_2.png\" alt=\"\"></p>\n<p>The first crop can't have equal paddings, so one side will have larger padding.<br>\n<img src=\"https://raw.githubusercontent.com/dimitreOliveira/MachineLearning/master/Kaggle/Rainforest%20Connection%20Species%20Audio%20Detection/Audio%20crop%20diagram.png\" alt=\"\"></p>\n<hr>\n<h3>The code</h3>\n<pre><code>def crop_audio(audio, tmin, tmax, min_crop_size=MIN_CROP_SIZE, sample_rate=48000, max_size=60.):\n    \"\"\"\n        Crops a 'waveform' file to have {min_crop_size} size given, {tmin}, {tmax}, {sample_rate} and {max_size}.\n    \"\"\"\n    label_size = tmax - tmin\n\n    if label_size &gt;= min_crop_size: # No padding needed\n        cut_min = tmin\n        cut_max = tmax        \n    else: # Needs padding\n        pad_start = (min_crop_size - label_size) / 2\n        cut_min = tf.maximum(0., (tmin - pad_start))\n\n        pad_end = (min_crop_size - (label_size - cut_min))\n        cut_max = tf.minimum(max_size, (tmax + pad_end))\n\n        cut_size = cut_max - cut_min\n\n        if cut_size &lt; min_crop_size:\n            cut_min = tf.maximum(0., (cut_min - (min_crop_size - cut_size)))\n\n    cut_size = cut_max - cut_min\n\n    # Casting tensors\n    cut_min = tf.cast((cut_min * sample_rate), tf.int32)\n    cut_max = tf.cast((cut_max * sample_rate), tf.int32)\n    cut_size = tf.cast((cut_size * sample_rate), tf.int32)\n\n    audio = audio[cut_min:cut_max] # croping the audio\n    audio = audio[:cut_size] # making sure it has the max size\n\n    audio = tf.reshape(audio, [cut_size]) # making sure it has the expected shape\n    return audio\n</code></pre>\n<pre><code>def random_crop_audio(audio, crop_size=CROP_SIZE, sample_rate=48000, max_size=60):\n    \"\"\"\n        Randomly crops a 'waveform' file to have {crop_size} size given, {sample_rate} and {max_size}.\n    \"\"\"\n    start = tf.random.uniform([], minval=0, \n                              maxval=(max_size - crop_size), \n                              dtype=tf.int32)\n    cut_min = start * sample_rate\n    cut_max = (start + crop_size) * sample_rate\n\n    audio_size = len(audio)\n    if cut_max &gt; audio_size:\n        cut_min -= cut_max - audio_size\n        cut_max = cut_min + (crop_size * sample_rate)\n\n    # Casting tensors\n    cut_min = tf.cast(cut_min, tf.int32)\n    cut_max = tf.cast(cut_max, tf.int32)\n    cut_size = tf.cast((crop_size*sample_rate), tf.int32)\n\n    audio = audio[cut_min:cut_max] # croping the audio\n    audio = audio[:cut_size] # making sure it has the max size\n\n    audio = tf.reshape(audio, [cut_size]) # making sure it has the expected shape\n    return audio\n</code></pre>\n<hr>\n<p>I am not very familiar with audio data but I will add a few observations here</p>\n<ul>\n<li>For me <code>EfficientNet</code> seems to work better than <code>ResNets</code>, I could not get ResNets to perform as well as public notebooks.</li>\n<li>Mel-spectrograms performed worse than regular spectrograms, but I might be doing something wrong.</li>\n<li>Tweaking the crop sizes seems more relevant than the spectrograms resolution (height x width)</li>\n<li>Incrementing the architecture at the top of the backbone matters very little.</li>\n</ul>",
      "rawMarkdown": "When I started looking at this competition one thing that confused me a little was how to effectively crop the audio, especially using Tensorflow to maximize the efficiency, I studied a few public notebooks and ended up making a lot of experiments into that, and have written a few functions that should be able to crop the data with some flexibility.\nBesides that, the importance of audio cropping has been discussed many times here, especially for creating an effective validation, references: [[1]](https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/200922), [[2]](https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/207624), [[3]](https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/215233), [[4]](https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/216564).\n\nThe Idea is to first crop the audio with a given `crop size` then apply **random crops** at this **first crop**. The idea is very similar to what most public notebooks uses, one thing that I tried to make sure of is that the first crop had **reasonable padding** in cases where the label was small.\n\nYou can find the notebook here: [Rainforest-Audio classification TF Improved](https://www.kaggle.com/dimitreoliveira/rainforest-audio-classification-tf-improved)\n\nTo help with the understanding let's look at these examples:\n\n**1st example**\nThe first crop can have equal paddings on both sides.\n![](https://raw.githubusercontent.com/dimitreOliveira/MachineLearning/master/Kaggle/Rainforest%20Connection%20Species%20Audio%20Detection/Audio%20crop%20diagram_2.png)\n\nThe first crop can't have equal paddings, so one side will have larger padding.\n![](https://raw.githubusercontent.com/dimitreOliveira/MachineLearning/master/Kaggle/Rainforest%20Connection%20Species%20Audio%20Detection/Audio%20crop%20diagram.png)\n\n---\n### The code\n\n```\ndef crop_audio(audio, tmin, tmax, min_crop_size=MIN_CROP_SIZE, sample_rate=48000, max_size=60.):\n    \"\"\"\n        Crops a 'waveform' file to have {min_crop_size} size given, {tmin}, {tmax}, {sample_rate} and {max_size}.\n    \"\"\"\n    label_size = tmax - tmin\n    \n    if label_size >= min_crop_size: # No padding needed\n        cut_min = tmin\n        cut_max = tmax        \n    else: # Needs padding\n        pad_start = (min_crop_size - label_size) / 2\n        cut_min = tf.maximum(0., (tmin - pad_start))\n        \n        pad_end = (min_crop_size - (label_size - cut_min))\n        cut_max = tf.minimum(max_size, (tmax + pad_end))\n        \n        cut_size = cut_max - cut_min\n        \n        if cut_size < min_crop_size:\n            cut_min = tf.maximum(0., (cut_min - (min_crop_size - cut_size)))\n        \n    cut_size = cut_max - cut_min\n    \n    # Casting tensors\n    cut_min = tf.cast((cut_min * sample_rate), tf.int32)\n    cut_max = tf.cast((cut_max * sample_rate), tf.int32)\n    cut_size = tf.cast((cut_size * sample_rate), tf.int32)\n    \n    audio = audio[cut_min:cut_max] # croping the audio\n    audio = audio[:cut_size] # making sure it has the max size\n    \n    audio = tf.reshape(audio, [cut_size]) # making sure it has the expected shape\n    return audio\n```\n```\ndef random_crop_audio(audio, crop_size=CROP_SIZE, sample_rate=48000, max_size=60):\n    \"\"\"\n        Randomly crops a 'waveform' file to have {crop_size} size given, {sample_rate} and {max_size}.\n    \"\"\"\n    start = tf.random.uniform([], minval=0, \n                              maxval=(max_size - crop_size), \n                              dtype=tf.int32)\n    cut_min = start * sample_rate\n    cut_max = (start + crop_size) * sample_rate\n    \n    audio_size = len(audio)\n    if cut_max > audio_size:\n        cut_min -= cut_max - audio_size\n        cut_max = cut_min + (crop_size * sample_rate)\n    \n    # Casting tensors\n    cut_min = tf.cast(cut_min, tf.int32)\n    cut_max = tf.cast(cut_max, tf.int32)\n    cut_size = tf.cast((crop_size*sample_rate), tf.int32)\n    \n    audio = audio[cut_min:cut_max] # croping the audio\n    audio = audio[:cut_size] # making sure it has the max size\n    \n    audio = tf.reshape(audio, [cut_size]) # making sure it has the expected shape\n    return audio\n```\n\n---\nI am not very familiar with audio data but I will add a few observations here\n\n- For me `EfficientNet` seems to work better than `ResNets`, I could not get ResNets to perform as well as public notebooks.\n- Mel-spectrograms performed worse than regular spectrograms, but I might be doing something wrong.\n- Tweaking the crop sizes seems more relevant than the spectrograms resolution (height x width)\n- Incrementing the architecture at the top of the backbone matters very little.",
      "votes": null
    },
    {
      "id": "1190917",
      "postDate": "02/08/2021 05:34:05",
      "content": "<p>Thanks for all your contributions Dimitre. Your notebooks are excellent. You share many good ideas.</p>",
      "rawMarkdown": "Thanks for all your contributions Dimitre. Your notebooks are excellent. You share many good ideas.",
      "votes": null
    },
    {
      "id": "1191313",
      "postDate": "02/08/2021 11:41:51",
      "content": "<p>Hey <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> , Thank you very much!</p>\n<p>Working with audio is very interesting, next time I will start earlier 😄</p>",
      "rawMarkdown": "Hey @cdeotte , Thank you very much!\n\nWorking with audio is very interesting, next time I will start earlier 😄",
      "votes": null
    },
    {
      "id": "1195517",
      "postDate": "02/10/2021 21:24:27",
      "content": "<p>Thanks for Sharing, I've been trying to crop audio with 2-4s length but don see any improved, maybe my model is not going well. Your notebook give more ideas, thanks again 👍</p>",
      "rawMarkdown": "Thanks for Sharing, I've been trying to crop audio with 2-4s length but don see any improved, maybe my model is not going well. Your notebook give more ideas, thanks again 👍",
      "votes": null
    },
    {
      "id": "1195526",
      "postDate": "02/10/2021 21:36:59",
      "content": "<p>You're welcome <a href=\"https://www.kaggle.com/scarecrow2020\" target=\"_blank\">@scarecrow2020</a> ,</p>\n<p>I did not have the TPU quota to experiment a lot, but I think that the best audio sizes should be somewhere around 4~10</p>",
      "rawMarkdown": "You're welcome @scarecrow2020 ,\n\nI did not have the TPU quota to experiment a lot, but I think that the best audio sizes should be somewhere around 4~10",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1190917,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "02/08/2021 05:34:05",
      "content": "<p>Thanks for all your contributions Dimitre. Your notebooks are excellent. You share many good ideas.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1191313,
          "author_name": "dimitreoliveira",
          "author_url": "",
          "post_date": "02/08/2021 11:41:51",
          "content": "<p>Hey <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> , Thank you very much!</p>\n<p>Working with audio is very interesting, next time I will start earlier 😄</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1195517,
      "author_name": "scarecrow2020",
      "author_url": "",
      "post_date": "02/10/2021 21:24:27",
      "content": "<p>Thanks for Sharing, I've been trying to crop audio with 2-4s length but don see any improved, maybe my model is not going well. Your notebook give more ideas, thanks again 👍</p>",
      "votes": null,
      "replies": [
        {
          "id": 1195526,
          "author_name": "dimitreoliveira",
          "author_url": "",
          "post_date": "02/10/2021 21:36:59",
          "content": "<p>You're welcome <a href=\"https://www.kaggle.com/scarecrow2020\" target=\"_blank\">@scarecrow2020</a> ,</p>\n<p>I did not have the TPU quota to experiment a lot, but I think that the best audio sizes should be somewhere around 4~10</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1189261": "When I started looking at this competition one thing that confused me a little was how to effectively crop the audio, especially using Tensorflow to maximize the efficiency, I studied a few public notebooks and ended up making a lot of experiments into that, and have written a few functions that should be able to crop the data with some flexibility.\nBesides that, the importance of audio cropping has been discussed many times here, especially for creating an effective validation, references: [[1]](https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/200922), [[2]](https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/207624), [[3]](https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/215233), [[4]](https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/216564).\n\nThe Idea is to first crop the audio with a given `crop size` then apply **random crops** at this **first crop**. The idea is very similar to what most public notebooks uses, one thing that I tried to make sure of is that the first crop had **reasonable padding** in cases where the label was small.\n\nYou can find the notebook here: [Rainforest-Audio classification TF Improved](https://www.kaggle.com/dimitreoliveira/rainforest-audio-classification-tf-improved)\n\nTo help with the understanding let's look at these examples:\n\n**1st example**\nThe first crop can have equal paddings on both sides.\n![](https://raw.githubusercontent.com/dimitreOliveira/MachineLearning/master/Kaggle/Rainforest%20Connection%20Species%20Audio%20Detection/Audio%20crop%20diagram_2.png)\n\nThe first crop can't have equal paddings, so one side will have larger padding.\n![](https://raw.githubusercontent.com/dimitreOliveira/MachineLearning/master/Kaggle/Rainforest%20Connection%20Species%20Audio%20Detection/Audio%20crop%20diagram.png)\n\n---\n### The code\n\n```\ndef crop_audio(audio, tmin, tmax, min_crop_size=MIN_CROP_SIZE, sample_rate=48000, max_size=60.):\n    \"\"\"\n        Crops a 'waveform' file to have {min_crop_size} size given, {tmin}, {tmax}, {sample_rate} and {max_size}.\n    \"\"\"\n    label_size = tmax - tmin\n    \n    if label_size >= min_crop_size: # No padding needed\n        cut_min = tmin\n        cut_max = tmax        \n    else: # Needs padding\n        pad_start = (min_crop_size - label_size) / 2\n        cut_min = tf.maximum(0., (tmin - pad_start))\n        \n        pad_end = (min_crop_size - (label_size - cut_min))\n        cut_max = tf.minimum(max_size, (tmax + pad_end))\n        \n        cut_size = cut_max - cut_min\n        \n        if cut_size < min_crop_size:\n            cut_min = tf.maximum(0., (cut_min - (min_crop_size - cut_size)))\n        \n    cut_size = cut_max - cut_min\n    \n    # Casting tensors\n    cut_min = tf.cast((cut_min * sample_rate), tf.int32)\n    cut_max = tf.cast((cut_max * sample_rate), tf.int32)\n    cut_size = tf.cast((cut_size * sample_rate), tf.int32)\n    \n    audio = audio[cut_min:cut_max] # croping the audio\n    audio = audio[:cut_size] # making sure it has the max size\n    \n    audio = tf.reshape(audio, [cut_size]) # making sure it has the expected shape\n    return audio\n```\n```\ndef random_crop_audio(audio, crop_size=CROP_SIZE, sample_rate=48000, max_size=60):\n    \"\"\"\n        Randomly crops a 'waveform' file to have {crop_size} size given, {sample_rate} and {max_size}.\n    \"\"\"\n    start = tf.random.uniform([], minval=0, \n                              maxval=(max_size - crop_size), \n                              dtype=tf.int32)\n    cut_min = start * sample_rate\n    cut_max = (start + crop_size) * sample_rate\n    \n    audio_size = len(audio)\n    if cut_max > audio_size:\n        cut_min -= cut_max - audio_size\n        cut_max = cut_min + (crop_size * sample_rate)\n    \n    # Casting tensors\n    cut_min = tf.cast(cut_min, tf.int32)\n    cut_max = tf.cast(cut_max, tf.int32)\n    cut_size = tf.cast((crop_size*sample_rate), tf.int32)\n    \n    audio = audio[cut_min:cut_max] # croping the audio\n    audio = audio[:cut_size] # making sure it has the max size\n    \n    audio = tf.reshape(audio, [cut_size]) # making sure it has the expected shape\n    return audio\n```\n\n---\nI am not very familiar with audio data but I will add a few observations here\n\n- For me `EfficientNet` seems to work better than `ResNets`, I could not get ResNets to perform as well as public notebooks.\n- Mel-spectrograms performed worse than regular spectrograms, but I might be doing something wrong.\n- Tweaking the crop sizes seems more relevant than the spectrograms resolution (height x width)\n- Incrementing the architecture at the top of the backbone matters very little.",
    "1190917": "Thanks for all your contributions Dimitre. Your notebooks are excellent. You share many good ideas.",
    "1191313": "Hey @cdeotte , Thank you very much!\n\nWorking with audio is very interesting, next time I will start earlier 😄",
    "1195517": "Thanks for Sharing, I've been trying to crop audio with 2-4s length but don see any improved, maybe my model is not going well. Your notebook give more ideas, thanks again 👍",
    "1195526": "You're welcome @scarecrow2020 ,\n\nI did not have the TPU quota to experiment a lot, but I think that the best audio sizes should be somewhere around 4~10"
  },
  "source": "meta"
}