{
  "id": 200922,
  "title": "Best way to crop audio",
  "url": "/competitions/rfcx-species-audio-detection/discussion/200922",
  "author_name": "Manh Lab",
  "post_date": "2020-12-02T12:07:39.970000",
  "votes": 23,
  "comment_count": 11,
  "views": 0,
  "content": "<p>I tried multi method to crop audio with difference length and start time. <br>\n10s cut with random start</p>\n<pre><code>        time_start = self.df.t_min.values[idx]*sr\n        time_end =   self.df.t_max.values[idx]*sr\n        # Positioning sound slice\n        center = np.round((time_start + time_end) / 2)\n        beginning = center - effective_length / 2\n        if beginning &lt; 0:\n            beginning = 0\n        beginning = np.random.randint( beginning , center)\n        ending = beginning + effective_length\n        if ending &gt; len(y):\n            ending = len(y)\n            beginning = ending - effective_length\n\n        y = y[beginning:ending].astype(np.float32)\n</code></pre>\n<p>from this notebook <a href=\"https://www.kaggle.com/fffrrt/all-in-one-rfcx-baseline-for-beginners\" target=\"_blank\">https://www.kaggle.com/fffrrt/all-in-one-rfcx-baseline-for-beginners</a></p>\n<pre><code>        time_start = self.df.t_min.values[idx]*sr\n        time_end =   self.df.t_max.values[idx]*sr\n        # Positioning sound slice\n        center = np.round((time_start + time_end) / 2)\n        beginning = center - effective_length / 2\n        if beginning &lt; 0:\n            beginning = 0\n        ending = beginning + effective_length\n        if ending &gt; len(y):\n            ending = len(y)\n            beginning = ending - effective_length\n\n        y = y[beginning:ending].astype(np.float32)\n</code></pre>\n<p>30s crop</p>\n<pre><code>            len_y = len(y)\n            effective_length = sr * PERIOD\n            if len_y &lt; effective_length:\n                new_y = np.zeros(effective_length, dtype=y.dtype)\n                start = np.random.randint(effective_length - len_y)\n                new_y[start:start + len_y] = y\n                y = new_y.astype(np.float32)\n            elif len_y &gt; effective_length:\n                start = np.random.randint(len_y - effective_length)\n                y = y[start:start + effective_length].astype(np.float32)\n            else:\n                y = y.astype(np.float32)\n</code></pre>",
  "messages": [
    {
      "id": 1102470,
      "postDate": "2020-12-05T01:17:55.720Z",
      "content": "<p>One thing you need to be careful is that you need to include all the labels in the cropped window. The code above is cropping around single event, but in the cropped clip there might be other tp events. Therefore, if you are to use the code above, label should be</p>\n<pre><code>time_start = self.df.t_min.values[idx]*sr\ntime_end = self.df.t_max.values[idx]*sr\n# Positioning sound slice\ncenter = np.round((time_start + time_end) / 2)\nbeginning = center - effective_length / 2\nif beginning &lt; 0:\n    beginning = 0\nbeginning = np.random.randint( beginning , center)\nending = beginning + effective_length\nif ending &gt; len(y):\n    ending = len(y)\nbeginning = ending - effective_length\ny = y[beginning:ending].astype(np.float32)\n\nbeginning_time = beginning / sr\nending_time = ending / sr\n\nrecording_id = self.df.loc[idx, \"recording_id\"]\nquery_string = f\"recording_id == '{recording_id}' &amp; \"\nquery_string += f\"t_min &lt; {ending_time} &amp; t_max &gt; {beginning_time}\"\nall_tp_events = self.df.query(query_string)\n\nlabel = np.zeros(24, dtype=np.float32)\nfor species_id in all_tp_events[\"species_id\"].unique():\n    label[int(species_id)] = 1.0\n</code></pre>",
      "rawMarkdown": "One thing you need to be careful is that you need to include all the labels in the cropped window. The code above is cropping around single event, but in the cropped clip there might be other tp events. Therefore, if you are to use the code above, label should be\n\n```\ntime_start = self.df.t_min.values[idx]*sr\ntime_end = self.df.t_max.values[idx]*sr\n# Positioning sound slice\ncenter = np.round((time_start + time_end) / 2)\nbeginning = center - effective_length / 2\nif beginning < 0:\n    beginning = 0\nbeginning = np.random.randint( beginning , center)\nending = beginning + effective_length\nif ending > len(y):\n    ending = len(y)\nbeginning = ending - effective_length\ny = y[beginning:ending].astype(np.float32)\n\nbeginning_time = beginning / sr\nending_time = ending / sr\n\nrecording_id = self.df.loc[idx, \"recording_id\"]\nquery_string = f\"recording_id == '{recording_id}' & \"\nquery_string += f\"t_min < {ending_time} & t_max > {beginning_time}\"\nall_tp_events = self.df.query(query_string)\n\nlabel = np.zeros(24, dtype=np.float32)\nfor species_id in all_tp_events[\"species_id\"].unique():\n    label[int(species_id)] = 1.0\n```",
      "votes": 33,
      "replies": [
        {
          "id": 1103517,
          "postDate": "2020-12-06T01:50:40.067Z",
          "content": "<p>yeah. And duplicated items need clear too.</p>",
          "rawMarkdown": "yeah. And duplicated items need clear too.",
          "votes": 1
        },
        {
          "id": 1151340,
          "postDate": "2021-01-13T08:52:16.947Z",
          "content": "<p>Thanks for sharing <a href=\"https://www.kaggle.com/hidehisaarai1213\" target=\"_blank\">@hidehisaarai1213</a> , may I ask you two questions for your code snippets:<br>\n1) Is it to iterate on rows of train_tp.csv, as idx is the index of each row?<br>\n2) What<code>s the purpose of np.random.randint(beginning , center) exactly? If I understand correctly, it</code>s in case for one flac/clip file has multiple event, when you iterate over train_tp.csv for this flac/clip file, you can generate different crop segment for this clip file, right?</p>",
          "rawMarkdown": "Thanks for sharing @hidehisaarai1213 , may I ask you two questions for your code snippets:\n1) Is it to iterate on rows of train_tp.csv, as idx is the index of each row?\n2) What`s the purpose of np.random.randint(beginning , center) exactly? If I understand correctly, it`s in case for one flac/clip file has multiple event, when you iterate over train_tp.csv for this flac/clip file, you can generate different crop segment for this clip file, right?"
        },
        {
          "id": 1152182,
          "postDate": "2021-01-13T22:21:13.510Z",
          "content": "<p>Hi,</p>\n<p>Interestingly enough, my score drops when I do the label correction as you stated at the end of your code. I have checked the labels and they seem correct. Do you have any ideas on that? What might be the issue apart from some weird bug in my code? </p>\n<p>Thanks 🙏</p>",
          "rawMarkdown": "Hi,\n\nInterestingly enough, my score drops when I do the label correction as you stated at the end of your code. I have checked the labels and they seem correct. Do you have any ideas on that? What might be the issue apart from some weird bug in my code? \n\nThanks 🙏",
          "votes": 1
        },
        {
          "id": 1154698,
          "postDate": "2021-01-15T19:45:39.990Z",
          "content": "<p>Same thing here. After reading this I realized that I don't include other labels in time window and I was looking for massive improvement in CV and LB, but nothing happened :D</p>",
          "rawMarkdown": "Same thing here. After reading this I realized that I don't include other labels in time window and I was looking for massive improvement in CV and LB, but nothing happened :D"
        }
      ]
    },
    {
      "id": 1099499,
      "postDate": "2020-12-02T12:07:39.970Z",
      "content": "<p>I tried multi method to crop audio with difference length and start time. <br>\n10s cut with random start</p>\n<pre><code>        time_start = self.df.t_min.values[idx]*sr\n        time_end =   self.df.t_max.values[idx]*sr\n        # Positioning sound slice\n        center = np.round((time_start + time_end) / 2)\n        beginning = center - effective_length / 2\n        if beginning &lt; 0:\n            beginning = 0\n        beginning = np.random.randint( beginning , center)\n        ending = beginning + effective_length\n        if ending &gt; len(y):\n            ending = len(y)\n            beginning = ending - effective_length\n\n        y = y[beginning:ending].astype(np.float32)\n</code></pre>\n<p>from this notebook <a href=\"https://www.kaggle.com/fffrrt/all-in-one-rfcx-baseline-for-beginners\" target=\"_blank\">https://www.kaggle.com/fffrrt/all-in-one-rfcx-baseline-for-beginners</a></p>\n<pre><code>        time_start = self.df.t_min.values[idx]*sr\n        time_end =   self.df.t_max.values[idx]*sr\n        # Positioning sound slice\n        center = np.round((time_start + time_end) / 2)\n        beginning = center - effective_length / 2\n        if beginning &lt; 0:\n            beginning = 0\n        ending = beginning + effective_length\n        if ending &gt; len(y):\n            ending = len(y)\n            beginning = ending - effective_length\n\n        y = y[beginning:ending].astype(np.float32)\n</code></pre>\n<p>30s crop</p>\n<pre><code>            len_y = len(y)\n            effective_length = sr * PERIOD\n            if len_y &lt; effective_length:\n                new_y = np.zeros(effective_length, dtype=y.dtype)\n                start = np.random.randint(effective_length - len_y)\n                new_y[start:start + len_y] = y\n                y = new_y.astype(np.float32)\n            elif len_y &gt; effective_length:\n                start = np.random.randint(len_y - effective_length)\n                y = y[start:start + effective_length].astype(np.float32)\n            else:\n                y = y.astype(np.float32)\n</code></pre>",
      "rawMarkdown": "I tried multi method to crop audio with difference length and start time. \n10s cut with random start\n```\n        time_start = self.df.t_min.values[idx]*sr\n        time_end =   self.df.t_max.values[idx]*sr\n        # Positioning sound slice\n        center = np.round((time_start + time_end) / 2)\n        beginning = center - effective_length / 2\n        if beginning < 0:\n            beginning = 0\n        beginning = np.random.randint( beginning , center)\n        ending = beginning + effective_length\n        if ending > len(y):\n            ending = len(y)\n            beginning = ending - effective_length\n\n        y = y[beginning:ending].astype(np.float32)\n```\nfrom this notebook https://www.kaggle.com/fffrrt/all-in-one-rfcx-baseline-for-beginners\n```\n        time_start = self.df.t_min.values[idx]*sr\n        time_end =   self.df.t_max.values[idx]*sr\n        # Positioning sound slice\n        center = np.round((time_start + time_end) / 2)\n        beginning = center - effective_length / 2\n        if beginning < 0:\n            beginning = 0\n        ending = beginning + effective_length\n        if ending > len(y):\n            ending = len(y)\n            beginning = ending - effective_length\n\n        y = y[beginning:ending].astype(np.float32)\n```\n30s crop\n```\n            len_y = len(y)\n            effective_length = sr * PERIOD\n            if len_y < effective_length:\n                new_y = np.zeros(effective_length, dtype=y.dtype)\n                start = np.random.randint(effective_length - len_y)\n                new_y[start:start + len_y] = y\n                y = new_y.astype(np.float32)\n            elif len_y > effective_length:\n                start = np.random.randint(len_y - effective_length)\n                y = y[start:start + effective_length].astype(np.float32)\n            else:\n                y = y.astype(np.float32)\n```\n",
      "votes": 23
    },
    {
      "id": 1099664,
      "postDate": "2020-12-02T14:13:53.883Z",
      "content": "<p>The first two methods don't crop or pad the signal with zeros which I think should be better as it will reduce edge effects when processing the signal slices with FFT, MelSpec etc…  </p>",
      "rawMarkdown": "The first two methods don't crop or pad the signal with zeros which I think should be better as it will reduce edge effects when processing the signal slices with FFT, MelSpec etc...  ",
      "votes": 1,
      "replies": [
        {
          "id": 1152356,
          "postDate": "2021-01-14T04:15:09.100Z",
          "content": "<p>Good idea, but this model performs poorly on inference. If you have some lifehacks - please share)</p>",
          "rawMarkdown": "Good idea, but this model performs poorly on inference. If you have some lifehacks - please share)",
          "votes": 2
        },
        {
          "id": 1178203,
          "postDate": "2021-01-30T18:05:04.107Z",
          "content": "<p><a href=\"https://www.kaggle.com/kupchanski\" target=\"_blank\">@kupchanski</a> Can you share some hacks pls</p>",
          "rawMarkdown": "@kupchanski Can you share some hacks pls"
        }
      ]
    },
    {
      "id": 1099678,
      "postDate": "2020-12-02T14:27:02.130Z",
      "rawMarkdown": "",
      "votes": 5,
      "isDeleted": true,
      "replies": [
        {
          "id": 1154817,
          "postDate": "2021-01-15T23:54:10.247Z",
          "content": "<p>Wouldn't it be simpler to calculate the spectogram for the entire audio clip and then index it? </p>",
          "rawMarkdown": "Wouldn't it be simpler to calculate the spectogram for the entire audio clip and then index it? "
        },
        {
          "id": 1157424,
          "postDate": "2021-01-17T22:09:45.383Z",
          "content": "<p>librosa uses hann windows when computing sttt, hence there is no need to worry about this.  If you use librosa that is.</p>",
          "rawMarkdown": "librosa uses hann windows when computing sttt, hence there is no need to worry about this.  If you use librosa that is."
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 1102470,
      "author_name": "Hidehisa Arai",
      "author_url": "",
      "post_date": "2020-12-05T01:17:55.720000",
      "content": "<p>One thing you need to be careful is that you need to include all the labels in the cropped window. The code above is cropping around single event, but in the cropped clip there might be other tp events. Therefore, if you are to use the code above, label should be</p>\n<pre><code>time_start = self.df.t_min.values[idx]*sr\ntime_end = self.df.t_max.values[idx]*sr\n# Positioning sound slice\ncenter = np.round((time_start + time_end) / 2)\nbeginning = center - effective_length / 2\nif beginning &lt; 0:\n    beginning = 0\nbeginning = np.random.randint( beginning , center)\nending = beginning + effective_length\nif ending &gt; len(y):\n    ending = len(y)\nbeginning = ending - effective_length\ny = y[beginning:ending].astype(np.float32)\n\nbeginning_time = beginning / sr\nending_time = ending / sr\n\nrecording_id = self.df.loc[idx, \"recording_id\"]\nquery_string = f\"recording_id == '{recording_id}' &amp; \"\nquery_string += f\"t_min &lt; {ending_time} &amp; t_max &gt; {beginning_time}\"\nall_tp_events = self.df.query(query_string)\n\nlabel = np.zeros(24, dtype=np.float32)\nfor species_id in all_tp_events[\"species_id\"].unique():\n    label[int(species_id)] = 1.0\n</code></pre>",
      "votes": 33,
      "replies": [
        {
          "id": 1103517,
          "author_name": "( ͡° ͜ʖ ͡°)",
          "author_url": "",
          "post_date": "2020-12-06T01:50:40.067000",
          "content": "<p>yeah. And duplicated items need clear too.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1151340,
          "author_name": "Hao",
          "author_url": "",
          "post_date": "2021-01-13T08:52:16.947000",
          "content": "<p>Thanks for sharing <a href=\"https://www.kaggle.com/hidehisaarai1213\" target=\"_blank\">@hidehisaarai1213</a> , may I ask you two questions for your code snippets:<br>\n1) Is it to iterate on rows of train_tp.csv, as idx is the index of each row?<br>\n2) What<code>s the purpose of np.random.randint(beginning , center) exactly? If I understand correctly, it</code>s in case for one flac/clip file has multiple event, when you iterate over train_tp.csv for this flac/clip file, you can generate different crop segment for this clip file, right?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1152182,
          "author_name": "Sinan Calisir",
          "author_url": "",
          "post_date": "2021-01-13T22:21:13.510000",
          "content": "<p>Hi,</p>\n<p>Interestingly enough, my score drops when I do the label correction as you stated at the end of your code. I have checked the labels and they seem correct. Do you have any ideas on that? What might be the issue apart from some weird bug in my code? </p>\n<p>Thanks 🙏</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1154698,
          "author_name": "Lubor Pacak",
          "author_url": "",
          "post_date": "2021-01-15T19:45:39.990000",
          "content": "<p>Same thing here. After reading this I realized that I don't include other labels in time window and I was looking for massive improvement in CV and LB, but nothing happened :D</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1099664,
      "author_name": "Jack Vial",
      "author_url": "",
      "post_date": "2020-12-02T14:13:53.883000",
      "content": "<p>The first two methods don't crop or pad the signal with zeros which I think should be better as it will reduce edge effects when processing the signal slices with FFT, MelSpec etc…  </p>",
      "votes": 1,
      "replies": [
        {
          "id": 1152356,
          "author_name": "Kupchanski",
          "author_url": "",
          "post_date": "2021-01-14T04:15:09.100000",
          "content": "<p>Good idea, but this model performs poorly on inference. If you have some lifehacks - please share)</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1178203,
          "author_name": "Gopi Durgaprasad",
          "author_url": "",
          "post_date": "2021-01-30T18:05:04.107000",
          "content": "<p><a href=\"https://www.kaggle.com/kupchanski\" target=\"_blank\">@kupchanski</a> Can you share some hacks pls</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1099678,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-12-02T14:27:02.130000",
      "content": "",
      "votes": 5,
      "replies": [
        {
          "id": 1154817,
          "author_name": "RNA",
          "author_url": "",
          "post_date": "2021-01-15T23:54:10.247000",
          "content": "<p>Wouldn't it be simpler to calculate the spectogram for the entire audio clip and then index it? </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1157424,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2021-01-17T22:09:45.383000",
          "content": "<p>librosa uses hann windows when computing sttt, hence there is no need to worry about this.  If you use librosa that is.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1102470": "One thing you need to be careful is that you need to include all the labels in the cropped window. The code above is cropping around single event, but in the cropped clip there might be other tp events. Therefore, if you are to use the code above, label should be\n\n```\ntime_start = self.df.t_min.values[idx]*sr\ntime_end = self.df.t_max.values[idx]*sr\n# Positioning sound slice\ncenter = np.round((time_start + time_end) / 2)\nbeginning = center - effective_length / 2\nif beginning < 0:\n    beginning = 0\nbeginning = np.random.randint( beginning , center)\nending = beginning + effective_length\nif ending > len(y):\n    ending = len(y)\nbeginning = ending - effective_length\ny = y[beginning:ending].astype(np.float32)\n\nbeginning_time = beginning / sr\nending_time = ending / sr\n\nrecording_id = self.df.loc[idx, \"recording_id\"]\nquery_string = f\"recording_id == '{recording_id}' & \"\nquery_string += f\"t_min < {ending_time} & t_max > {beginning_time}\"\nall_tp_events = self.df.query(query_string)\n\nlabel = np.zeros(24, dtype=np.float32)\nfor species_id in all_tp_events[\"species_id\"].unique():\n    label[int(species_id)] = 1.0\n```",
    "1099499": "I tried multi method to crop audio with difference length and start time. \n10s cut with random start\n```\n        time_start = self.df.t_min.values[idx]*sr\n        time_end =   self.df.t_max.values[idx]*sr\n        # Positioning sound slice\n        center = np.round((time_start + time_end) / 2)\n        beginning = center - effective_length / 2\n        if beginning < 0:\n            beginning = 0\n        beginning = np.random.randint( beginning , center)\n        ending = beginning + effective_length\n        if ending > len(y):\n            ending = len(y)\n            beginning = ending - effective_length\n\n        y = y[beginning:ending].astype(np.float32)\n```\nfrom this notebook https://www.kaggle.com/fffrrt/all-in-one-rfcx-baseline-for-beginners\n```\n        time_start = self.df.t_min.values[idx]*sr\n        time_end =   self.df.t_max.values[idx]*sr\n        # Positioning sound slice\n        center = np.round((time_start + time_end) / 2)\n        beginning = center - effective_length / 2\n        if beginning < 0:\n            beginning = 0\n        ending = beginning + effective_length\n        if ending > len(y):\n            ending = len(y)\n            beginning = ending - effective_length\n\n        y = y[beginning:ending].astype(np.float32)\n```\n30s crop\n```\n            len_y = len(y)\n            effective_length = sr * PERIOD\n            if len_y < effective_length:\n                new_y = np.zeros(effective_length, dtype=y.dtype)\n                start = np.random.randint(effective_length - len_y)\n                new_y[start:start + len_y] = y\n                y = new_y.astype(np.float32)\n            elif len_y > effective_length:\n                start = np.random.randint(len_y - effective_length)\n                y = y[start:start + effective_length].astype(np.float32)\n            else:\n                y = y.astype(np.float32)\n```\n",
    "1099664": "The first two methods don't crop or pad the signal with zeros which I think should be better as it will reduce edge effects when processing the signal slices with FFT, MelSpec etc...  ",
    "1099678": ""
  }
}