{
  "id": 243671,
  "title": "Part of the 18th place solution (Naoism's model)",
  "url": "/competitions/birdclef-2021/discussion/243671",
  "author_name": "",
  "post_date": "2021-06-03T14:36:05.312641200Z",
  "votes": 13,
  "comment_count": 4,
  "views": 0,
  "content": "<p>At first, I would like to express my thanks to all the organizers for hosting this interesting competition and my teammates.<br>\nAnd thanks to all other competitors for sharing valuable information in forums and notebooks. </p>\n<p>My goal was to win a gold medal in this competition and become a master. But I couldn't. I will try my best to become a master in the next competition.</p>\n<h2>The challenges of this competition</h2>\n<p>(1) Label weakness and noise,<br>\n(2) Domain mismatch between train and test data</p>\n<h2>Summary</h2>\n<ul>\n<li>I referred to <a href=\"https://www.kaggle.com/theoviel\" target=\"_blank\">@theoviel</a> 's solution in the previous birdcall competition. (<a href=\"https://www.kaggle.com/c/birdsong-recognition/discussion/183199\" target=\"_blank\">https://www.kaggle.com/c/birdsong-recognition/discussion/183199</a>) <br>\nI would like to say thanks here. Thank you for sharing valuable solution and code. </li>\n<li>Model algorithm is Resnest50d</li>\n<li>The model was trained on train short audios (only audio length &lt;= 60) (Approach to Challenge (1))</li>\n<li>Augmentation: Gaussian noise, Background noise, Mixup (Approach to Challenge (2))</li>\n<li>I labeled some of the data nocall time range and \"human voice\" time range manually. (Approach to Challenge (1))</li>\n<li>2 type postprocessing: <ul>\n<li>Ensemble the prediction results of the 2.5s shift sequence. (Proposed by <a href=\"https://www.kaggle.com/Iafoss\" target=\"_blank\">@Iafoss</a>)</li>\n<li>Multiply all chunk predictions for a given class by 1.3 or 1.35 if the predicted value for the given class over the entire sequence</li></ul></li>\n<li>My best single 5fold model's score is 0.6206(Private) /0.7471(Public)</li>\n</ul>\n<h2>Training data</h2>\n<ul>\n<li><p>Improved randomly crop the read data into 5s chunk.</p>\n<ul>\n<li>Based on the assumption that there are often birds at the beginning and end of the audio (confirmed by EDA), crop the beginning of the audio (randomly changed from 0-2s) with a probability of 20%, the end of the audio (randomly changed from 0-2s) with a probability of 20%, and completely randomly with a probability of 60%.</li></ul></li>\n<li><p>Train dataset: train short audios data with an audio length of 60 seconds or less</p>\n<ul>\n<li>The longer the audio length, the higher the probability of cropping the nocall position or the position of the bird without the label on it.</li>\n<li>Reduced the number of data and training time.</li></ul></li>\n</ul>\n<h2>Data Augmentation</h2>\n<ul>\n<li><p>Gaussian noise</p></li>\n<li><p>Background noise</p>\n<ul>\n<li>Just like the previous competition, in this competition we also had to deal with domain shifts. So I trained my models by adding the following noises to the background.</li>\n<li>The nocall part from the example_test_audio in Cornell Birdcall Identification in 2020. <a href=\"https://www.kaggle.com/c/birdsong-recognition/overview\" target=\"_blank\">https://www.kaggle.com/c/birdsong-recognition/overview</a></li>\n<li>Data of Rainforest Connection Species Audio Detection (<a href=\"https://www.kaggle.com/c/rfcx-species-audio-detection/overview\" target=\"_blank\">https://www.kaggle.com/c/rfcx-species-audio-detection/overview</a>)<br>\nThis is <a href=\"https://www.kaggle.com/Iafoss\" target=\"_blank\">@Iafoss</a> 's idea. Thanks.</li>\n<li>Pink noise</li></ul></li>\n<li><p>Mixup</p>\n<ul>\n<li>I expanded the Modified Mixup(proposed by <a href=\"https://www.kaggle.com/theoviel\" target=\"_blank\">@theoviel</a>) to 3audio and used.</li>\n<li>I applied the Modified Mixup with 50% probability, 2audio with 80% probability, and 3audio with 20% probability.</li></ul></li>\n</ul>\n<h2>Hand labeling</h2>\n<ul>\n<li><p>Nocall</p>\n<ul>\n<li>Hand labeling is time consuming. So, to save time, I did not try to find the location of the birds, but just noted the time when the nocall range was large.</li>\n<li>And there is also a risk. The risk is that there are too many classes to distinguish, and thus the positions of the primary and secondary labels may get mixed up. So, I did this only for audio that did not have a secondary label and was less than 16 seconds long. These audios are unlikely to contain more than one bird. Therefore this helps to reduce that risk.</li></ul></li>\n<li><p>Human voice</p>\n<ul>\n<li>What is human voice ?</li>\n<li>Some audios contain digitized human voices, not environmental sounds or even bird sounds. This is displayed in the spectrogram as follows. When I thought about why such audio was included, I assumed that it came from a particular author who put it there to edit the audio. So I checked the author data in the human voice data I found during the nocall hand labeling and EDA. As a result, this assumption turned out to be correct.  So, I visualized the author's data and labeled the human voice part. When loading the data in the model training, we cut off that part before using it.<br>\nEx.1) row_ids: whiwre1/XC518092.ogg<br>\nEx.2) row_ids: willet1/XC148079.ogg<br>\n<img src=\"https://user-images.githubusercontent.com/58644957/120661891-a15fa980-c4c3-11eb-9447-5c355a048a18.png\" alt=\"image (1)\"><br>\n<img src=\"https://user-images.githubusercontent.com/58644957/120661953-b0def280-c4c3-11eb-8907-0db20d8e285b.png\" alt=\"image (2)\"></li></ul></li>\n</ul>\n<h2>Masked loss</h2>\n<ul>\n<li>The primary labels are noisy, but the secondary labels are even noisier. Therefore, I masked the loss of the secondary labels only when doing mixup. This idea was based on <a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> 's solution in the previous competition. Thank you for sharing useful ideas. <a href=\"https://www.kaggle.com/c/birdsong-recognition/discussion/183219\" target=\"_blank\">https://www.kaggle.com/c/birdsong-recognition/discussion/183219</a></li>\n</ul>\n<h2>Ensemble</h2>\n<ul>\n<li>Instead of a simple average, I used powed average(p=2 or p=4)</li>\n</ul>\n<h2>Post processing</h2>\n<ul>\n<li>The test data was recorded at the same location for 10 minutes. Therefore, the types of birds that appear have a higher probability of appearing in the same recording.<br>\nSo, I thought of a post-processing to slightly increase the prediction probability (I multiply the prediction probability by 1.3 ~ 1.35) of birds predicted in the same recording.</li>\n<li>Ensemble the prediction results of the 2.5s shift sequence. (Proposed by <a href=\"https://www.kaggle.com/Iafoss\" target=\"_blank\">@Iafoss</a>)</li>\n</ul>",
  "messages": [
    {
      "id": "1334482",
      "postDate": "06/03/2021 14:36:05",
      "content": "<p>At first, I would like to express my thanks to all the organizers for hosting this interesting competition and my teammates.<br>\nAnd thanks to all other competitors for sharing valuable information in forums and notebooks. </p>\n<p>My goal was to win a gold medal in this competition and become a master. But I couldn't. I will try my best to become a master in the next competition.</p>\n<h2>The challenges of this competition</h2>\n<p>(1) Label weakness and noise,<br>\n(2) Domain mismatch between train and test data</p>\n<h2>Summary</h2>\n<ul>\n<li>I referred to <a href=\"https://www.kaggle.com/theoviel\" target=\"_blank\">@theoviel</a> 's solution in the previous birdcall competition. (<a href=\"https://www.kaggle.com/c/birdsong-recognition/discussion/183199\" target=\"_blank\">https://www.kaggle.com/c/birdsong-recognition/discussion/183199</a>) <br>\nI would like to say thanks here. Thank you for sharing valuable solution and code. </li>\n<li>Model algorithm is Resnest50d</li>\n<li>The model was trained on train short audios (only audio length &lt;= 60) (Approach to Challenge (1))</li>\n<li>Augmentation: Gaussian noise, Background noise, Mixup (Approach to Challenge (2))</li>\n<li>I labeled some of the data nocall time range and \"human voice\" time range manually. (Approach to Challenge (1))</li>\n<li>2 type postprocessing: <ul>\n<li>Ensemble the prediction results of the 2.5s shift sequence. (Proposed by <a href=\"https://www.kaggle.com/Iafoss\" target=\"_blank\">@Iafoss</a>)</li>\n<li>Multiply all chunk predictions for a given class by 1.3 or 1.35 if the predicted value for the given class over the entire sequence</li></ul></li>\n<li>My best single 5fold model's score is 0.6206(Private) /0.7471(Public)</li>\n</ul>\n<h2>Training data</h2>\n<ul>\n<li><p>Improved randomly crop the read data into 5s chunk.</p>\n<ul>\n<li>Based on the assumption that there are often birds at the beginning and end of the audio (confirmed by EDA), crop the beginning of the audio (randomly changed from 0-2s) with a probability of 20%, the end of the audio (randomly changed from 0-2s) with a probability of 20%, and completely randomly with a probability of 60%.</li></ul></li>\n<li><p>Train dataset: train short audios data with an audio length of 60 seconds or less</p>\n<ul>\n<li>The longer the audio length, the higher the probability of cropping the nocall position or the position of the bird without the label on it.</li>\n<li>Reduced the number of data and training time.</li></ul></li>\n</ul>\n<h2>Data Augmentation</h2>\n<ul>\n<li><p>Gaussian noise</p></li>\n<li><p>Background noise</p>\n<ul>\n<li>Just like the previous competition, in this competition we also had to deal with domain shifts. So I trained my models by adding the following noises to the background.</li>\n<li>The nocall part from the example_test_audio in Cornell Birdcall Identification in 2020. <a href=\"https://www.kaggle.com/c/birdsong-recognition/overview\" target=\"_blank\">https://www.kaggle.com/c/birdsong-recognition/overview</a></li>\n<li>Data of Rainforest Connection Species Audio Detection (<a href=\"https://www.kaggle.com/c/rfcx-species-audio-detection/overview\" target=\"_blank\">https://www.kaggle.com/c/rfcx-species-audio-detection/overview</a>)<br>\nThis is <a href=\"https://www.kaggle.com/Iafoss\" target=\"_blank\">@Iafoss</a> 's idea. Thanks.</li>\n<li>Pink noise</li></ul></li>\n<li><p>Mixup</p>\n<ul>\n<li>I expanded the Modified Mixup(proposed by <a href=\"https://www.kaggle.com/theoviel\" target=\"_blank\">@theoviel</a>) to 3audio and used.</li>\n<li>I applied the Modified Mixup with 50% probability, 2audio with 80% probability, and 3audio with 20% probability.</li></ul></li>\n</ul>\n<h2>Hand labeling</h2>\n<ul>\n<li><p>Nocall</p>\n<ul>\n<li>Hand labeling is time consuming. So, to save time, I did not try to find the location of the birds, but just noted the time when the nocall range was large.</li>\n<li>And there is also a risk. The risk is that there are too many classes to distinguish, and thus the positions of the primary and secondary labels may get mixed up. So, I did this only for audio that did not have a secondary label and was less than 16 seconds long. These audios are unlikely to contain more than one bird. Therefore this helps to reduce that risk.</li></ul></li>\n<li><p>Human voice</p>\n<ul>\n<li>What is human voice ?</li>\n<li>Some audios contain digitized human voices, not environmental sounds or even bird sounds. This is displayed in the spectrogram as follows. When I thought about why such audio was included, I assumed that it came from a particular author who put it there to edit the audio. So I checked the author data in the human voice data I found during the nocall hand labeling and EDA. As a result, this assumption turned out to be correct.  So, I visualized the author's data and labeled the human voice part. When loading the data in the model training, we cut off that part before using it.<br>\nEx.1) row_ids: whiwre1/XC518092.ogg<br>\nEx.2) row_ids: willet1/XC148079.ogg<br>\n<img src=\"https://user-images.githubusercontent.com/58644957/120661891-a15fa980-c4c3-11eb-9447-5c355a048a18.png\" alt=\"image (1)\"><br>\n<img src=\"https://user-images.githubusercontent.com/58644957/120661953-b0def280-c4c3-11eb-8907-0db20d8e285b.png\" alt=\"image (2)\"></li></ul></li>\n</ul>\n<h2>Masked loss</h2>\n<ul>\n<li>The primary labels are noisy, but the secondary labels are even noisier. Therefore, I masked the loss of the secondary labels only when doing mixup. This idea was based on <a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> 's solution in the previous competition. Thank you for sharing useful ideas. <a href=\"https://www.kaggle.com/c/birdsong-recognition/discussion/183219\" target=\"_blank\">https://www.kaggle.com/c/birdsong-recognition/discussion/183219</a></li>\n</ul>\n<h2>Ensemble</h2>\n<ul>\n<li>Instead of a simple average, I used powed average(p=2 or p=4)</li>\n</ul>\n<h2>Post processing</h2>\n<ul>\n<li>The test data was recorded at the same location for 10 minutes. Therefore, the types of birds that appear have a higher probability of appearing in the same recording.<br>\nSo, I thought of a post-processing to slightly increase the prediction probability (I multiply the prediction probability by 1.3 ~ 1.35) of birds predicted in the same recording.</li>\n<li>Ensemble the prediction results of the 2.5s shift sequence. (Proposed by <a href=\"https://www.kaggle.com/Iafoss\" target=\"_blank\">@Iafoss</a>)</li>\n</ul>",
      "rawMarkdown": "At first, I would like to express my thanks to all the organizers for hosting this interesting competition and my teammates.\nAnd thanks to all other competitors for sharing valuable information in forums and notebooks. \n\nMy goal was to win a gold medal in this competition and become a master. But I couldn't. I will try my best to become a master in the next competition.\n\n## The challenges of this competition\n(1) Label weakness and noise,\n(2) Domain mismatch between train and test data\n\n## Summary\n- I referred to @theoviel 's solution in the previous birdcall competition. (https://www.kaggle.com/c/birdsong-recognition/discussion/183199) \nI would like to say thanks here. Thank you for sharing valuable solution and code. \n- Model algorithm is Resnest50d\n- The model was trained on train short audios (only audio length <= 60) (Approach to Challenge (1))\n- Augmentation: Gaussian noise, Background noise, Mixup (Approach to Challenge (2))\n- I labeled some of the data nocall time range and \"human voice\" time range manually. (Approach to Challenge (1))\n- 2 type postprocessing: \n  - Ensemble the prediction results of the 2.5s shift sequence. (Proposed by @Iafoss)\n  - Multiply all chunk predictions for a given class by 1.3 or 1.35 if the predicted value for the given class over the entire sequence\n- My best single 5fold model's score is 0.6206(Private) /0.7471(Public)\n\n\n## Training data\n- Improved randomly crop the read data into 5s chunk.\n  - Based on the assumption that there are often birds at the beginning and end of the audio (confirmed by EDA), crop the beginning of the audio (randomly changed from 0-2s) with a probability of 20%, the end of the audio (randomly changed from 0-2s) with a probability of 20%, and completely randomly with a probability of 60%.\n\n- Train dataset: train short audios data with an audio length of 60 seconds or less\n  - The longer the audio length, the higher the probability of cropping the nocall position or the position of the bird without the label on it.\n  - Reduced the number of data and training time.\n\n\n## Data Augmentation\n\n- Gaussian noise\n- Background noise\n  - Just like the previous competition, in this competition we also had to deal with domain shifts. So I trained my models by adding the following noises to the background.\n  - The nocall part from the example_test_audio in Cornell Birdcall Identification in 2020. https://www.kaggle.com/c/birdsong-recognition/overview\n  - Data of Rainforest Connection Species Audio Detection (https://www.kaggle.com/c/rfcx-species-audio-detection/overview)\n  This is @Iafoss 's idea. Thanks.\n  - Pink noise\n\n- Mixup\n  - I expanded the Modified Mixup(proposed by @theoviel) to 3audio and used.\n  - I applied the Modified Mixup with 50% probability, 2audio with 80% probability, and 3audio with 20% probability.\n\n## Hand labeling\n\n- Nocall\n  - Hand labeling is time consuming. So, to save time, I did not try to find the location of the birds, but just noted the time when the nocall range was large.\n  - And there is also a risk. The risk is that there are too many classes to distinguish, and thus the positions of the primary and secondary labels may get mixed up. So, I did this only for audio that did not have a secondary label and was less than 16 seconds long. These audios are unlikely to contain more than one bird. Therefore this helps to reduce that risk.\n\n- Human voice\n  - What is human voice ?\n  - Some audios contain digitized human voices, not environmental sounds or even bird sounds. This is displayed in the spectrogram as follows. When I thought about why such audio was included, I assumed that it came from a particular author who put it there to edit the audio. So I checked the author data in the human voice data I found during the nocall hand labeling and EDA. As a result, this assumption turned out to be correct.  So, I visualized the author's data and labeled the human voice part. When loading the data in the model training, we cut off that part before using it.\n    Ex.1) row_ids: whiwre1/XC518092.ogg\n    Ex.2) row_ids: willet1/XC148079.ogg\n![image (1)](https://user-images.githubusercontent.com/58644957/120661891-a15fa980-c4c3-11eb-9447-5c355a048a18.png)\n![image (2)](https://user-images.githubusercontent.com/58644957/120661953-b0def280-c4c3-11eb-8907-0db20d8e285b.png)\n\n## Masked loss\n  - The primary labels are noisy, but the secondary labels are even noisier. Therefore, I masked the loss of the secondary labels only when doing mixup. This idea was based on @cpmpml 's solution in the previous competition. Thank you for sharing useful ideas. https://www.kaggle.com/c/birdsong-recognition/discussion/183219\n\n## Ensemble\n  - Instead of a simple average, I used powed average(p=2 or p=4)\n\n## Post processing\n  - The test data was recorded at the same location for 10 minutes. Therefore, the types of birds that appear have a higher probability of appearing in the same recording.\n  So, I thought of a post-processing to slightly increase the prediction probability (I multiply the prediction probability by 1.3 ~ 1.35) of birds predicted in the same recording.\n  - Ensemble the prediction results of the 2.5s shift sequence. (Proposed by @Iafoss)",
      "votes": null
    },
    {
      "id": "1334541",
      "postDate": "06/03/2021 15:25:37",
      "content": "<p>I'm glad my code helped you, and congratz on the nice finish !</p>",
      "rawMarkdown": "I'm glad my code helped you, and congratz on the nice finish !",
      "votes": null
    },
    {
      "id": "1334583",
      "postDate": "06/03/2021 15:55:53",
      "content": "<p>Congrats on the end result, and thanks for the citation!</p>",
      "rawMarkdown": "Congrats on the end result, and thanks for the citation!",
      "votes": null
    },
    {
      "id": "1334986",
      "postDate": "06/04/2021 00:18:33",
      "content": "<p>Thanks! Your solution helped my team a lot.</p>",
      "rawMarkdown": "Thanks! Your solution helped my team a lot.",
      "votes": null
    },
    {
      "id": "1334987",
      "postDate": "06/04/2021 00:18:49",
      "content": "<p>Thanks!        </p>",
      "rawMarkdown": "Thanks!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1334541,
      "author_name": "theoviel",
      "author_url": "",
      "post_date": "06/03/2021 15:25:37",
      "content": "<p>I'm glad my code helped you, and congratz on the nice finish !</p>",
      "votes": null,
      "replies": [
        {
          "id": 1334986,
          "author_name": "naoism",
          "author_url": "",
          "post_date": "06/04/2021 00:18:33",
          "content": "<p>Thanks! Your solution helped my team a lot.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1334583,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "06/03/2021 15:55:53",
      "content": "<p>Congrats on the end result, and thanks for the citation!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1334987,
          "author_name": "naoism",
          "author_url": "",
          "post_date": "06/04/2021 00:18:49",
          "content": "<p>Thanks!        </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1334482": "At first, I would like to express my thanks to all the organizers for hosting this interesting competition and my teammates.\nAnd thanks to all other competitors for sharing valuable information in forums and notebooks. \n\nMy goal was to win a gold medal in this competition and become a master. But I couldn't. I will try my best to become a master in the next competition.\n\n## The challenges of this competition\n(1) Label weakness and noise,\n(2) Domain mismatch between train and test data\n\n## Summary\n- I referred to @theoviel 's solution in the previous birdcall competition. (https://www.kaggle.com/c/birdsong-recognition/discussion/183199) \nI would like to say thanks here. Thank you for sharing valuable solution and code. \n- Model algorithm is Resnest50d\n- The model was trained on train short audios (only audio length <= 60) (Approach to Challenge (1))\n- Augmentation: Gaussian noise, Background noise, Mixup (Approach to Challenge (2))\n- I labeled some of the data nocall time range and \"human voice\" time range manually. (Approach to Challenge (1))\n- 2 type postprocessing: \n  - Ensemble the prediction results of the 2.5s shift sequence. (Proposed by @Iafoss)\n  - Multiply all chunk predictions for a given class by 1.3 or 1.35 if the predicted value for the given class over the entire sequence\n- My best single 5fold model's score is 0.6206(Private) /0.7471(Public)\n\n\n## Training data\n- Improved randomly crop the read data into 5s chunk.\n  - Based on the assumption that there are often birds at the beginning and end of the audio (confirmed by EDA), crop the beginning of the audio (randomly changed from 0-2s) with a probability of 20%, the end of the audio (randomly changed from 0-2s) with a probability of 20%, and completely randomly with a probability of 60%.\n\n- Train dataset: train short audios data with an audio length of 60 seconds or less\n  - The longer the audio length, the higher the probability of cropping the nocall position or the position of the bird without the label on it.\n  - Reduced the number of data and training time.\n\n\n## Data Augmentation\n\n- Gaussian noise\n- Background noise\n  - Just like the previous competition, in this competition we also had to deal with domain shifts. So I trained my models by adding the following noises to the background.\n  - The nocall part from the example_test_audio in Cornell Birdcall Identification in 2020. https://www.kaggle.com/c/birdsong-recognition/overview\n  - Data of Rainforest Connection Species Audio Detection (https://www.kaggle.com/c/rfcx-species-audio-detection/overview)\n  This is @Iafoss 's idea. Thanks.\n  - Pink noise\n\n- Mixup\n  - I expanded the Modified Mixup(proposed by @theoviel) to 3audio and used.\n  - I applied the Modified Mixup with 50% probability, 2audio with 80% probability, and 3audio with 20% probability.\n\n## Hand labeling\n\n- Nocall\n  - Hand labeling is time consuming. So, to save time, I did not try to find the location of the birds, but just noted the time when the nocall range was large.\n  - And there is also a risk. The risk is that there are too many classes to distinguish, and thus the positions of the primary and secondary labels may get mixed up. So, I did this only for audio that did not have a secondary label and was less than 16 seconds long. These audios are unlikely to contain more than one bird. Therefore this helps to reduce that risk.\n\n- Human voice\n  - What is human voice ?\n  - Some audios contain digitized human voices, not environmental sounds or even bird sounds. This is displayed in the spectrogram as follows. When I thought about why such audio was included, I assumed that it came from a particular author who put it there to edit the audio. So I checked the author data in the human voice data I found during the nocall hand labeling and EDA. As a result, this assumption turned out to be correct.  So, I visualized the author's data and labeled the human voice part. When loading the data in the model training, we cut off that part before using it.\n    Ex.1) row_ids: whiwre1/XC518092.ogg\n    Ex.2) row_ids: willet1/XC148079.ogg\n![image (1)](https://user-images.githubusercontent.com/58644957/120661891-a15fa980-c4c3-11eb-9447-5c355a048a18.png)\n![image (2)](https://user-images.githubusercontent.com/58644957/120661953-b0def280-c4c3-11eb-8907-0db20d8e285b.png)\n\n## Masked loss\n  - The primary labels are noisy, but the secondary labels are even noisier. Therefore, I masked the loss of the secondary labels only when doing mixup. This idea was based on @cpmpml 's solution in the previous competition. Thank you for sharing useful ideas. https://www.kaggle.com/c/birdsong-recognition/discussion/183219\n\n## Ensemble\n  - Instead of a simple average, I used powed average(p=2 or p=4)\n\n## Post processing\n  - The test data was recorded at the same location for 10 minutes. Therefore, the types of birds that appear have a higher probability of appearing in the same recording.\n  So, I thought of a post-processing to slightly increase the prediction probability (I multiply the prediction probability by 1.3 ~ 1.35) of birds predicted in the same recording.\n  - Ensemble the prediction results of the 2.5s shift sequence. (Proposed by @Iafoss)",
    "1334541": "I'm glad my code helped you, and congratz on the nice finish !",
    "1334583": "Congrats on the end result, and thanks for the citation!",
    "1334986": "Thanks! Your solution helped my team a lot.",
    "1334987": "Thanks!"
  },
  "source": "meta"
}