{
  "id": 201827,
  "title": "Test window size for mel specs",
  "url": "/competitions/rfcx-species-audio-detection/discussion/201827",
  "author_name": "",
  "post_date": "2020-12-07T01:09:49.102598300Z",
  "votes": 5,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Hi guys! </p>\n<p>I trained my NN on Mel Specs from 10 sec cropped audios. I was wondering which should be the optimal window size for testing. I was thinking taking the whole audio, create a Mel Spec and re-size them to the same size I used for training.</p>\n<p>Which strategy are you using/do you think is the best?</p>",
  "messages": [
    {
      "id": "1104463",
      "postDate": "12/07/2020 01:09:49",
      "content": "<p>Hi guys! </p>\n<p>I trained my NN on Mel Specs from 10 sec cropped audios. I was wondering which should be the optimal window size for testing. I was thinking taking the whole audio, create a Mel Spec and re-size them to the same size I used for training.</p>\n<p>Which strategy are you using/do you think is the best?</p>",
      "rawMarkdown": "Hi guys! \n\nI trained my NN on Mel Specs from 10 sec cropped audios. I was wondering which should be the optimal window size for testing. I was thinking taking the whole audio, create a Mel Spec and re-size them to the same size I used for training.\n\nWhich strategy are you using/do you think is the best?",
      "votes": null
    },
    {
      "id": "1106585",
      "postDate": "12/08/2020 23:58:39",
      "content": "<p>Training on 10 second clips and testing on 60 second clips won't be the most accurate approach, because the 60 second clip will be squeezed to fit the same size (that is, the \"width\" of the spectrogram image). Your test data will therefore look skinnier than your train data, and will have some trouble classifying.</p>\n<p>A different approach would be to make multiple predictions per test audio clip. For example, maybe make 12 predictions, each of 10 seconds, with 5 seconds of overlap. From there, take the maximum predicted value for each class. </p>\n<p>Taking the maximum is important, since not every 10s test clip will contain the bird. You could also do some more sophisticated techniques for aggregating your prediction values, but just taking the max is a good start.</p>\n<p>Good luck!</p>",
      "rawMarkdown": "Training on 10 second clips and testing on 60 second clips won't be the most accurate approach, because the 60 second clip will be squeezed to fit the same size (that is, the \"width\" of the spectrogram image). Your test data will therefore look skinnier than your train data, and will have some trouble classifying.\n\nA different approach would be to make multiple predictions per test audio clip. For example, maybe make 12 predictions, each of 10 seconds, with 5 seconds of overlap. From there, take the maximum predicted value for each class. \n\nTaking the maximum is important, since not every 10s test clip will contain the bird. You could also do some more sophisticated techniques for aggregating your prediction values, but just taking the max is a good start.\n\nGood luck!",
      "votes": null
    },
    {
      "id": "1106612",
      "postDate": "12/09/2020 00:38:01",
      "content": "<p>Thanks for your response Mike. I was thinking the same about the first approach. Regarding training with several 10 second clips I think might take a lot of time though as computing 1992*6 is quite expensive, not to mention that testing data is only 21%. </p>\n<p>I look forward to any other approaches other competitors might be using to tackle this issue. </p>\n<p>Thanks again!</p>",
      "rawMarkdown": "Thanks for your response Mike. I was thinking the same about the first approach. Regarding training with several 10 second clips I think might take a lot of time though as computing 1992*6 is quite expensive, not to mention that testing data is only 21%. \n\nI look forward to any other approaches other competitors might be using to tackle this issue. \n\nThanks again!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1106585,
      "author_name": "maltonji",
      "author_url": "",
      "post_date": "12/08/2020 23:58:39",
      "content": "<p>Training on 10 second clips and testing on 60 second clips won't be the most accurate approach, because the 60 second clip will be squeezed to fit the same size (that is, the \"width\" of the spectrogram image). Your test data will therefore look skinnier than your train data, and will have some trouble classifying.</p>\n<p>A different approach would be to make multiple predictions per test audio clip. For example, maybe make 12 predictions, each of 10 seconds, with 5 seconds of overlap. From there, take the maximum predicted value for each class. </p>\n<p>Taking the maximum is important, since not every 10s test clip will contain the bird. You could also do some more sophisticated techniques for aggregating your prediction values, but just taking the max is a good start.</p>\n<p>Good luck!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1106612,
          "author_name": "alejopaullier",
          "author_url": "",
          "post_date": "12/09/2020 00:38:01",
          "content": "<p>Thanks for your response Mike. I was thinking the same about the first approach. Regarding training with several 10 second clips I think might take a lot of time though as computing 1992*6 is quite expensive, not to mention that testing data is only 21%. </p>\n<p>I look forward to any other approaches other competitors might be using to tackle this issue. </p>\n<p>Thanks again!</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1104463": "Hi guys! \n\nI trained my NN on Mel Specs from 10 sec cropped audios. I was wondering which should be the optimal window size for testing. I was thinking taking the whole audio, create a Mel Spec and re-size them to the same size I used for training.\n\nWhich strategy are you using/do you think is the best?",
    "1106585": "Training on 10 second clips and testing on 60 second clips won't be the most accurate approach, because the 60 second clip will be squeezed to fit the same size (that is, the \"width\" of the spectrogram image). Your test data will therefore look skinnier than your train data, and will have some trouble classifying.\n\nA different approach would be to make multiple predictions per test audio clip. For example, maybe make 12 predictions, each of 10 seconds, with 5 seconds of overlap. From there, take the maximum predicted value for each class. \n\nTaking the maximum is important, since not every 10s test clip will contain the bird. You could also do some more sophisticated techniques for aggregating your prediction values, but just taking the max is a good start.\n\nGood luck!",
    "1106612": "Thanks for your response Mike. I was thinking the same about the first approach. Regarding training with several 10 second clips I think might take a lot of time though as computing 1992*6 is quite expensive, not to mention that testing data is only 21%. \n\nI look forward to any other approaches other competitors might be using to tackle this issue. \n\nThanks again!"
  },
  "source": "meta"
}