{
  "id": 313068,
  "title": "🤔 Train on 7 Second Clips but Inference on 5 Second Clips?",
  "url": "/competitions/birdclef-2022/discussion/313068",
  "author_name": "",
  "post_date": "2022-03-15T13:29:22.099143800Z",
  "votes": 6,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Hi there, I've noticed this phenomenon in the past BirdCLEF competitions and it will probably occur in this competition as well.</p>\n<p>The main example of this technique are <a href=\"https://www.kaggle.com/kneroma\" target=\"_blank\"><strong>kneroma's</strong></a> notebooks for <a href=\"https://www.kaggle.com/kneroma/clean-fast-simple-bird-identifier-training-colab\" target=\"_blank\"><strong>training</strong></a> and <a href=\"https://www.kaggle.com/kneroma/clean-fast-simple-bird-identifier-inference\" target=\"_blank\"><strong>inference</strong></a>.</p>\n<p>The thing I've noticed is:</p>\n<ul>\n<li>Train is done on 7-second clips (128,281,3) MEL spectrograms. </li>\n<li>However, we perform inference on 5-second clips (128,201,3) MEL spectrograms.</li>\n</ul>\n<p>Would this not stretch the spectrograms? I've noticed in some of the published papers, the reason given is that training on 7-second segments 'works better' and is 'more likely to contain bird calls'. I like the idea of replicating previous competition solutions, but I like to understand why they did what they did. In this case, I'm most of the way there…. but I don't understand this inference part? Can anyone (or <a href=\"https://www.kaggle.com/kneroma\" target=\"_blank\">@kneroma</a> ?) clarify the logic here and why this doesn't result in worse performance?</p>\n<hr>\n<p>Just as an aside… I think I may handle it by using some sort of overlapping segment system.</p>\n<ul>\n<li>i.e. Pad beginning and end of long audio clip w/ 6 seconds of blank audio</li>\n<li>Split the test audio clip into 7-second clips with a 1-second step</li>\n<li>Use a weighted average (based on the percentage of overlap between the 7-second and 5-second segment) of all predictions.<ul>\n<li>i.e. when a 7-second segment completely overlaps a 5-second segment, it will have a weight of 1 and those predictions will contribute strongly. When a 7-second segment only overlaps with HALF a 5-second segment, its weight will be 0.5. </li></ul></li>\n</ul>\n<p><br></p>\n<p>I took the time to create a diagram illustrating what this inference would look like for a single 5-second test clip given a step size of 2 instead of 1 (mostly because it's easier to draw). <em>Note the weighting is wrong for the 3/5 overlap… duh… obviously this should have a weight of 0.6 not 0.5…</em></p>\n<p><img src=\"https://i.ibb.co/6gQy9d3/Inferenec-Weighting-drawio-1.png\" alt=\"single_clip\"></p>\n<hr>\n<p>If anyone can offer clarity on my above question or comment on this potential inference approach, it would be greatly appreciated. Thanks in advance!</p>",
  "messages": [
    {
      "id": "1723505",
      "postDate": "03/15/2022 13:29:22",
      "content": "<p>Hi there, I've noticed this phenomenon in the past BirdCLEF competitions and it will probably occur in this competition as well.</p>\n<p>The main example of this technique are <a href=\"https://www.kaggle.com/kneroma\" target=\"_blank\"><strong>kneroma's</strong></a> notebooks for <a href=\"https://www.kaggle.com/kneroma/clean-fast-simple-bird-identifier-training-colab\" target=\"_blank\"><strong>training</strong></a> and <a href=\"https://www.kaggle.com/kneroma/clean-fast-simple-bird-identifier-inference\" target=\"_blank\"><strong>inference</strong></a>.</p>\n<p>The thing I've noticed is:</p>\n<ul>\n<li>Train is done on 7-second clips (128,281,3) MEL spectrograms. </li>\n<li>However, we perform inference on 5-second clips (128,201,3) MEL spectrograms.</li>\n</ul>\n<p>Would this not stretch the spectrograms? I've noticed in some of the published papers, the reason given is that training on 7-second segments 'works better' and is 'more likely to contain bird calls'. I like the idea of replicating previous competition solutions, but I like to understand why they did what they did. In this case, I'm most of the way there…. but I don't understand this inference part? Can anyone (or <a href=\"https://www.kaggle.com/kneroma\" target=\"_blank\">@kneroma</a> ?) clarify the logic here and why this doesn't result in worse performance?</p>\n<hr>\n<p>Just as an aside… I think I may handle it by using some sort of overlapping segment system.</p>\n<ul>\n<li>i.e. Pad beginning and end of long audio clip w/ 6 seconds of blank audio</li>\n<li>Split the test audio clip into 7-second clips with a 1-second step</li>\n<li>Use a weighted average (based on the percentage of overlap between the 7-second and 5-second segment) of all predictions.<ul>\n<li>i.e. when a 7-second segment completely overlaps a 5-second segment, it will have a weight of 1 and those predictions will contribute strongly. When a 7-second segment only overlaps with HALF a 5-second segment, its weight will be 0.5. </li></ul></li>\n</ul>\n<p><br></p>\n<p>I took the time to create a diagram illustrating what this inference would look like for a single 5-second test clip given a step size of 2 instead of 1 (mostly because it's easier to draw). <em>Note the weighting is wrong for the 3/5 overlap… duh… obviously this should have a weight of 0.6 not 0.5…</em></p>\n<p><img src=\"https://i.ibb.co/6gQy9d3/Inferenec-Weighting-drawio-1.png\" alt=\"single_clip\"></p>\n<hr>\n<p>If anyone can offer clarity on my above question or comment on this potential inference approach, it would be greatly appreciated. Thanks in advance!</p>",
      "rawMarkdown": "Hi there, I've noticed this phenomenon in the past BirdCLEF competitions and it will probably occur in this competition as well.\n\nThe main example of this technique are [**kneroma's**](https://www.kaggle.com/kneroma) notebooks for [**training**](https://www.kaggle.com/kneroma/clean-fast-simple-bird-identifier-training-colab) and [**inference**](https://www.kaggle.com/kneroma/clean-fast-simple-bird-identifier-inference).\n\nThe thing I've noticed is:\n* Train is done on 7-second clips (128,281,3) MEL spectrograms. \n* However, we perform inference on 5-second clips (128,201,3) MEL spectrograms.\n\nWould this not stretch the spectrograms? I've noticed in some of the published papers, the reason given is that training on 7-second segments 'works better' and is 'more likely to contain bird calls'. I like the idea of replicating previous competition solutions, but I like to understand why they did what they did. In this case, I'm most of the way there.... but I don't understand this inference part? Can anyone (or @kneroma ?) clarify the logic here and why this doesn't result in worse performance?\n\n---\n\nJust as an aside... I think I may handle it by using some sort of overlapping segment system.\n* i.e. Pad beginning and end of long audio clip w/ 6 seconds of blank audio\n* Split the test audio clip into 7-second clips with a 1-second step\n* Use a weighted average (based on the percentage of overlap between the 7-second and 5-second segment) of all predictions.\n  * i.e. when a 7-second segment completely overlaps a 5-second segment, it will have a weight of 1 and those predictions will contribute strongly. When a 7-second segment only overlaps with HALF a 5-second segment, its weight will be 0.5. \n\n<br>\n\nI took the time to create a diagram illustrating what this inference would look like for a single 5-second test clip given a step size of 2 instead of 1 (mostly because it's easier to draw). *Note the weighting is wrong for the 3/5 overlap... duh... obviously this should have a weight of 0.6 not 0.5...*\n\n![single_clip](https://i.ibb.co/6gQy9d3/Inferenec-Weighting-drawio-1.png)\n\n---\n\nIf anyone can offer clarity on my above question or comment on this potential inference approach, it would be greatly appreciated. Thanks in advance!",
      "votes": null
    },
    {
      "id": "1723550",
      "postDate": "03/15/2022 14:09:24",
      "content": "<p>I wish I could give you an answer right now but not only i'm tired but also I have to read again these old works to be sure I'm not telling you nonsense ! Good competition and hope I can join before the end since audio competitions are always interesting ones.</p>",
      "rawMarkdown": "I wish I could give you an answer right now but not only i'm tired but also I have to read again these old works to be sure I'm not telling you nonsense ! Good competition and hope I can join before the end since audio competitions are always interesting ones.",
      "votes": null
    },
    {
      "id": "1723571",
      "postDate": "03/15/2022 14:29:05",
      "content": "<p>Looking forward to it! :)</p>",
      "rawMarkdown": "Looking forward to it! :)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1723550,
      "author_name": "kneroma",
      "author_url": "",
      "post_date": "03/15/2022 14:09:24",
      "content": "<p>I wish I could give you an answer right now but not only i'm tired but also I have to read again these old works to be sure I'm not telling you nonsense ! Good competition and hope I can join before the end since audio competitions are always interesting ones.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1723571,
          "author_name": "dschettler8845",
          "author_url": "",
          "post_date": "03/15/2022 14:29:05",
          "content": "<p>Looking forward to it! :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1723505": "Hi there, I've noticed this phenomenon in the past BirdCLEF competitions and it will probably occur in this competition as well.\n\nThe main example of this technique are [**kneroma's**](https://www.kaggle.com/kneroma) notebooks for [**training**](https://www.kaggle.com/kneroma/clean-fast-simple-bird-identifier-training-colab) and [**inference**](https://www.kaggle.com/kneroma/clean-fast-simple-bird-identifier-inference).\n\nThe thing I've noticed is:\n* Train is done on 7-second clips (128,281,3) MEL spectrograms. \n* However, we perform inference on 5-second clips (128,201,3) MEL spectrograms.\n\nWould this not stretch the spectrograms? I've noticed in some of the published papers, the reason given is that training on 7-second segments 'works better' and is 'more likely to contain bird calls'. I like the idea of replicating previous competition solutions, but I like to understand why they did what they did. In this case, I'm most of the way there.... but I don't understand this inference part? Can anyone (or @kneroma ?) clarify the logic here and why this doesn't result in worse performance?\n\n---\n\nJust as an aside... I think I may handle it by using some sort of overlapping segment system.\n* i.e. Pad beginning and end of long audio clip w/ 6 seconds of blank audio\n* Split the test audio clip into 7-second clips with a 1-second step\n* Use a weighted average (based on the percentage of overlap between the 7-second and 5-second segment) of all predictions.\n  * i.e. when a 7-second segment completely overlaps a 5-second segment, it will have a weight of 1 and those predictions will contribute strongly. When a 7-second segment only overlaps with HALF a 5-second segment, its weight will be 0.5. \n\n<br>\n\nI took the time to create a diagram illustrating what this inference would look like for a single 5-second test clip given a step size of 2 instead of 1 (mostly because it's easier to draw). *Note the weighting is wrong for the 3/5 overlap... duh... obviously this should have a weight of 0.6 not 0.5...*\n\n![single_clip](https://i.ibb.co/6gQy9d3/Inferenec-Weighting-drawio-1.png)\n\n---\n\nIf anyone can offer clarity on my above question or comment on this potential inference approach, it would be greatly appreciated. Thanks in advance!",
    "1723550": "I wish I could give you an answer right now but not only i'm tired but also I have to read again these old works to be sure I'm not telling you nonsense ! Good competition and hope I can join before the end since audio competitions are always interesting ones.",
    "1723571": "Looking forward to it! :)"
  },
  "source": "meta"
}