{
  "id": 202217,
  "title": "Set Window Size per Class",
  "url": "/competitions/rfcx-species-audio-detection/discussion/202217",
  "author_name": "",
  "post_date": "2020-12-09T00:21:45.109948800Z",
  "votes": 5,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Most people in the notebooks and discussions are using a 10s window size for training. I think this is because the maximum time range is ~8s, so 10s sufficiently captures it.</p>\n<p>However, most of the class time ranges are under 3 seconds (some of which are 0.25s!). If a class is detected in a 0.25 second window, but the clip is 10 seconds, wouldn't that be hard to find? It's like trying to detect a tiny object in a large image.</p>\n<p>I'm considering setting the window size separately for each class, then after computing all the different spectrograms, resizing the images to be the same before training the model. This means for each audio file, many spectrograms will need to be produced (ones with a 1s window, others with a 3s window, etc.). Additionally, it'll make inference more complex. That said, I think it'd yield some benefits for predicting the classes with smaller time ranges. Any thoughts?</p>",
  "messages": [
    {
      "id": "1106605",
      "postDate": "12/09/2020 00:21:45",
      "content": "<p>Most people in the notebooks and discussions are using a 10s window size for training. I think this is because the maximum time range is ~8s, so 10s sufficiently captures it.</p>\n<p>However, most of the class time ranges are under 3 seconds (some of which are 0.25s!). If a class is detected in a 0.25 second window, but the clip is 10 seconds, wouldn't that be hard to find? It's like trying to detect a tiny object in a large image.</p>\n<p>I'm considering setting the window size separately for each class, then after computing all the different spectrograms, resizing the images to be the same before training the model. This means for each audio file, many spectrograms will need to be produced (ones with a 1s window, others with a 3s window, etc.). Additionally, it'll make inference more complex. That said, I think it'd yield some benefits for predicting the classes with smaller time ranges. Any thoughts?</p>",
      "rawMarkdown": "Most people in the notebooks and discussions are using a 10s window size for training. I think this is because the maximum time range is ~8s, so 10s sufficiently captures it.\n\nHowever, most of the class time ranges are under 3 seconds (some of which are 0.25s!). If a class is detected in a 0.25 second window, but the clip is 10 seconds, wouldn't that be hard to find? It's like trying to detect a tiny object in a large image.\n\nI'm considering setting the window size separately for each class, then after computing all the different spectrograms, resizing the images to be the same before training the model. This means for each audio file, many spectrograms will need to be produced (ones with a 1s window, others with a 3s window, etc.). Additionally, it'll make inference more complex. That said, I think it'd yield some benefits for predicting the classes with smaller time ranges. Any thoughts?",
      "votes": null
    },
    {
      "id": "1109374",
      "postDate": "12/11/2020 15:58:13",
      "content": "<p>But with this different time windows, you want to have only one NN model to the 24 classes or will you do one NN model for each of the 24 classes and join the results?</p>\n<p>I think that the problem of using only one NN for the 24 classes is in the prediction, how would you cut the test audio for predictions? </p>",
      "rawMarkdown": "But with this different time windows, you want to have only one NN model to the 24 classes or will you do one NN model for each of the 24 classes and join the results?\n\nI think that the problem of using only one NN for the 24 classes is in the prediction, how would you cut the test audio for predictions?",
      "votes": null
    },
    {
      "id": "1109441",
      "postDate": "12/11/2020 17:23:11",
      "content": "<p>I think I'd find middle-ground, where maybe there are 4 models:</p>\n<ul>\n<li>High/Low time range</li>\n<li>High/Low frequency<br>\n(1 model would be high time range, high frequency. Another would be high time range, low frequency. etc.)</li>\n</ul>\n<p>Since I'd have 4 models, I would cut the test audio 4 different ways described above. Does that answer your question?</p>",
      "rawMarkdown": "I think I'd find middle-ground, where maybe there are 4 models:\n- High/Low time range\n- High/Low frequency\n(1 model would be high time range, high frequency. Another would be high time range, low frequency. etc.)\n\nSince I'd have 4 models, I would cut the test audio 4 different ways described above. Does that answer your question?",
      "votes": null
    },
    {
      "id": "1109448",
      "postDate": "12/11/2020 17:32:15",
      "content": "<p>Of course, I'm going in the same direction. I was thinking about making 24 NN models, but I think it would be too much to train and too slow to make the inference. I am analyzing a sweet spot on how many models to use.</p>",
      "rawMarkdown": "Of course, I'm going in the same direction. I was thinking about making 24 NN models, but I think it would be too much to train and too slow to make the inference. I am analyzing a sweet spot on how many models to use.",
      "votes": null
    },
    {
      "id": "1109475",
      "postDate": "12/11/2020 18:20:26",
      "content": "<p>Agreed! Since this isn't a notebook competition, inference speed doesn't matter. Obviously we'd hope it doesn't take <strong>too long</strong> to make predictions, but at least we don't have to worry about a 4 hour time limit or anything.</p>",
      "rawMarkdown": "Agreed! Since this isn't a notebook competition, inference speed doesn't matter. Obviously we'd hope it doesn't take **too long** to make predictions, but at least we don't have to worry about a 4 hour time limit or anything.",
      "votes": null
    },
    {
      "id": "1109494",
      "postDate": "12/11/2020 18:47:32",
      "content": "<p>Yes, but the problem with inference speed is for testing new ideas and parameters. But there's a long way until the end of this competition! haha</p>",
      "rawMarkdown": "Yes, but the problem with inference speed is for testing new ideas and parameters. But there's a long way until the end of this competition! haha",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1109374,
      "author_name": "willsoares1",
      "author_url": "",
      "post_date": "12/11/2020 15:58:13",
      "content": "<p>But with this different time windows, you want to have only one NN model to the 24 classes or will you do one NN model for each of the 24 classes and join the results?</p>\n<p>I think that the problem of using only one NN for the 24 classes is in the prediction, how would you cut the test audio for predictions? </p>",
      "votes": null,
      "replies": [
        {
          "id": 1109441,
          "author_name": "maltonji",
          "author_url": "",
          "post_date": "12/11/2020 17:23:11",
          "content": "<p>I think I'd find middle-ground, where maybe there are 4 models:</p>\n<ul>\n<li>High/Low time range</li>\n<li>High/Low frequency<br>\n(1 model would be high time range, high frequency. Another would be high time range, low frequency. etc.)</li>\n</ul>\n<p>Since I'd have 4 models, I would cut the test audio 4 different ways described above. Does that answer your question?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1109448,
          "author_name": "willsoares1",
          "author_url": "",
          "post_date": "12/11/2020 17:32:15",
          "content": "<p>Of course, I'm going in the same direction. I was thinking about making 24 NN models, but I think it would be too much to train and too slow to make the inference. I am analyzing a sweet spot on how many models to use.</p>",
          "votes": null,
          "replies": [
            {
              "id": 1109475,
              "author_name": "maltonji",
              "author_url": "",
              "post_date": "12/11/2020 18:20:26",
              "content": "<p>Agreed! Since this isn't a notebook competition, inference speed doesn't matter. Obviously we'd hope it doesn't take <strong>too long</strong> to make predictions, but at least we don't have to worry about a 4 hour time limit or anything.</p>",
              "votes": null,
              "replies": []
            }
          ]
        },
        {
          "id": 1109494,
          "author_name": "willsoares1",
          "author_url": "",
          "post_date": "12/11/2020 18:47:32",
          "content": "<p>Yes, but the problem with inference speed is for testing new ideas and parameters. But there's a long way until the end of this competition! haha</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1106605": "Most people in the notebooks and discussions are using a 10s window size for training. I think this is because the maximum time range is ~8s, so 10s sufficiently captures it.\n\nHowever, most of the class time ranges are under 3 seconds (some of which are 0.25s!). If a class is detected in a 0.25 second window, but the clip is 10 seconds, wouldn't that be hard to find? It's like trying to detect a tiny object in a large image.\n\nI'm considering setting the window size separately for each class, then after computing all the different spectrograms, resizing the images to be the same before training the model. This means for each audio file, many spectrograms will need to be produced (ones with a 1s window, others with a 3s window, etc.). Additionally, it'll make inference more complex. That said, I think it'd yield some benefits for predicting the classes with smaller time ranges. Any thoughts?",
    "1109374": "But with this different time windows, you want to have only one NN model to the 24 classes or will you do one NN model for each of the 24 classes and join the results?\n\nI think that the problem of using only one NN for the 24 classes is in the prediction, how would you cut the test audio for predictions?",
    "1109441": "I think I'd find middle-ground, where maybe there are 4 models:\n- High/Low time range\n- High/Low frequency\n(1 model would be high time range, high frequency. Another would be high time range, low frequency. etc.)\n\nSince I'd have 4 models, I would cut the test audio 4 different ways described above. Does that answer your question?",
    "1109448": "Of course, I'm going in the same direction. I was thinking about making 24 NN models, but I think it would be too much to train and too slow to make the inference. I am analyzing a sweet spot on how many models to use.",
    "1109475": "Agreed! Since this isn't a notebook competition, inference speed doesn't matter. Obviously we'd hope it doesn't take **too long** to make predictions, but at least we don't have to worry about a 4 hour time limit or anything.",
    "1109494": "Yes, but the problem with inference speed is for testing new ideas and parameters. But there's a long way until the end of this competition! haha"
  },
  "source": "meta"
}