{
  "id": 243351,
  "title": "5th place solution",
  "url": "/competitions/birdclef-2021/writeups/kramarenko-vladislav-5th-place-solution",
  "author_name": "",
  "post_date": "2021-06-02T06:48:28.929109300Z",
  "votes": 38,
  "comment_count": 8,
  "views": 0,
  "content": "<p>Thanks to the organizers for an interesting contest and <a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> for the motivation.</p>\n<p><strong>My decision is based on my public code posted here:</strong><br>\n<a href=\"https://www.kaggle.com/c/birdsong-recognition/discussion/183269\" target=\"_blank\">https://www.kaggle.com/c/birdsong-recognition/discussion/183269</a><br>\n<a href=\"https://github.com/vlomme/Birdcall-Identification-competition\" target=\"_blank\">https://github.com/vlomme/Birdcall-Identification-competition</a></p>\n<p><strong>The basic solution easily gets silver. What I changed:</strong></p>\n<ul>\n<li>Switched to SED +1%.</li>\n<li>Lowered threshold +1% . </li>\n<li>Left only birds from region of record +0%</li>\n<li>Changed ensemble averaging +1%</li>\n</ul>\n<h2>Pre-processing:</h2>\n<p>Used log-melspectrograms<br>\nn_fft = 1536, sr = 21952, hop_length = 245, n_mels = 224, len_chack 448, image_size = 224 * 448, 5 seconds</p>\n<h2>Models:</h2>\n<p>Ensemble of 14 models (sed_resnet50, sed_resnest50, sed _efficientnet-b0)</p>\n<h2>Augmentations:</h2>\n<ul>\n<li>For contrast, I raised the image to a power of 0.5 to 3. at 0.5, the background noise is closer to the birds, and at 3, on the contrary, the quiet sounds become even quieter.</li>\n<li>Slightly accelerated / slowed down recording</li>\n<li>Add a different sound without birds(rain, noise, conversations, etc.)</li>\n<li>Added white, pink, and band noise. Increasing the noise level increases recall, but reduces precision.</li>\n<li>With a probability of 0.5 lowered the upper frequencies. In the real world, the upper frequencies fade faster with distance</li>\n</ul>\n<h2>Train:</h2>\n<ul>\n<li>Used BCEWithLogitsLoss. For the main birds, the label was 1. For birds in the background 0.3.</li>\n<li>Used loss:</li>\n</ul>\n<pre><code>train_los1 = nn.BCEWithLogitsLoss()(prediction['clipwise_output'], true)\ntrain_los2 = nn.BCEWithLogitsLoss()(prediction[\"segmentwise_output_max\"], true)\ntrain_loss = (train_los1 + train_los2)/2\n</code></pre>\n<ul>\n<li>I didn't look at metrics on training records, but only on validation files (train_soundscapes)</li>\n<li>20-40 epochs (24 hours on gtx1060)</li>\n</ul>\n<h2>Postprocessing</h2>\n<ul>\n<li>If there was a bird in the segment, I increased the probability of finding it in the entire file.</li>\n<li>Model ensemble averaging</li>\n</ul>\n<pre><code>proba1 = proba.prod(axis = 0) ** (1.0/len(proba))\nproba = proba**2\nproba = proba.mean(axis=0)\nproba = proba**(1/2)\nproba = (proba + proba1)/2\n</code></pre>\n<p>Since the training took a long time and other approaches didn't work, I switched to other competitions and have hardly taught any new models in the last month. It's a shame that a little was not enough, I will try to do better</p>",
  "messages": [
    {
      "id": "1332514",
      "postDate": "06/02/2021 06:48:28",
      "content": "<p>Thanks to the organizers for an interesting contest and <a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> for the motivation.</p>\n<p><strong>My decision is based on my public code posted here:</strong><br>\n<a href=\"https://www.kaggle.com/c/birdsong-recognition/discussion/183269\" target=\"_blank\">https://www.kaggle.com/c/birdsong-recognition/discussion/183269</a><br>\n<a href=\"https://github.com/vlomme/Birdcall-Identification-competition\" target=\"_blank\">https://github.com/vlomme/Birdcall-Identification-competition</a></p>\n<p><strong>The basic solution easily gets silver. What I changed:</strong></p>\n<ul>\n<li>Switched to SED +1%.</li>\n<li>Lowered threshold +1% . </li>\n<li>Left only birds from region of record +0%</li>\n<li>Changed ensemble averaging +1%</li>\n</ul>\n<h2>Pre-processing:</h2>\n<p>Used log-melspectrograms<br>\nn_fft = 1536, sr = 21952, hop_length = 245, n_mels = 224, len_chack 448, image_size = 224 * 448, 5 seconds</p>\n<h2>Models:</h2>\n<p>Ensemble of 14 models (sed_resnet50, sed_resnest50, sed _efficientnet-b0)</p>\n<h2>Augmentations:</h2>\n<ul>\n<li>For contrast, I raised the image to a power of 0.5 to 3. at 0.5, the background noise is closer to the birds, and at 3, on the contrary, the quiet sounds become even quieter.</li>\n<li>Slightly accelerated / slowed down recording</li>\n<li>Add a different sound without birds(rain, noise, conversations, etc.)</li>\n<li>Added white, pink, and band noise. Increasing the noise level increases recall, but reduces precision.</li>\n<li>With a probability of 0.5 lowered the upper frequencies. In the real world, the upper frequencies fade faster with distance</li>\n</ul>\n<h2>Train:</h2>\n<ul>\n<li>Used BCEWithLogitsLoss. For the main birds, the label was 1. For birds in the background 0.3.</li>\n<li>Used loss:</li>\n</ul>\n<pre><code>train_los1 = nn.BCEWithLogitsLoss()(prediction['clipwise_output'], true)\ntrain_los2 = nn.BCEWithLogitsLoss()(prediction[\"segmentwise_output_max\"], true)\ntrain_loss = (train_los1 + train_los2)/2\n</code></pre>\n<ul>\n<li>I didn't look at metrics on training records, but only on validation files (train_soundscapes)</li>\n<li>20-40 epochs (24 hours on gtx1060)</li>\n</ul>\n<h2>Postprocessing</h2>\n<ul>\n<li>If there was a bird in the segment, I increased the probability of finding it in the entire file.</li>\n<li>Model ensemble averaging</li>\n</ul>\n<pre><code>proba1 = proba.prod(axis = 0) ** (1.0/len(proba))\nproba = proba**2\nproba = proba.mean(axis=0)\nproba = proba**(1/2)\nproba = (proba + proba1)/2\n</code></pre>\n<p>Since the training took a long time and other approaches didn't work, I switched to other competitions and have hardly taught any new models in the last month. It's a shame that a little was not enough, I will try to do better</p>",
      "rawMarkdown": "Thanks to the organizers for an interesting contest and @cpmpml for the motivation.\n\n**My decision is based on my public code posted here:**\nhttps://www.kaggle.com/c/birdsong-recognition/discussion/183269\nhttps://github.com/vlomme/Birdcall-Identification-competition\n\n**The basic solution easily gets silver. What I changed:**\n- Switched to SED +1%.\n- Lowered threshold +1% . \n- Left only birds from region of record +0%\n- Changed ensemble averaging +1%\n\n## Pre-processing:\nUsed log-melspectrograms\nn_fft = 1536, sr = 21952, hop_length = 245, n_mels = 224, len_chack 448, image_size = 224 * 448, 5 seconds\n\n## Models:\nEnsemble of 14 models (sed_resnet50, sed_resnest50, sed _efficientnet-b0)\n\n## Augmentations:\n- For contrast, I raised the image to a power of 0.5 to 3. at 0.5, the background noise is closer to the birds, and at 3, on the contrary, the quiet sounds become even quieter.\n- Slightly accelerated / slowed down recording\n- Add a different sound without birds(rain, noise, conversations, etc.)\n- Added white, pink, and band noise. Increasing the noise level increases recall, but reduces precision.\n- With a probability of 0.5 lowered the upper frequencies. In the real world, the upper frequencies fade faster with distance\n\n## Train:\n- Used BCEWithLogitsLoss. For the main birds, the label was 1. For birds in the background 0.3.\n- Used loss:\n```\ntrain_los1 = nn.BCEWithLogitsLoss()(prediction['clipwise_output'], true)\ntrain_los2 = nn.BCEWithLogitsLoss()(prediction[\"segmentwise_output_max\"], true)\ntrain_loss = (train_los1 + train_los2)/2\n```\n- I didn't look at metrics on training records, but only on validation files (train_soundscapes)\n- 20-40 epochs (24 hours on gtx1060)\n\n## Postprocessing\n- If there was a bird in the segment, I increased the probability of finding it in the entire file.\n- Model ensemble averaging\n```\nproba1 = proba.prod(axis = 0) ** (1.0/len(proba))\nproba = proba**2\nproba = proba.mean(axis=0)\nproba = proba**(1/2)\nproba = (proba + proba1)/2\n```\n\nSince the training took a long time and other approaches didn't work, I switched to other competitions and have hardly taught any new models in the last month. It's a shame that a little was not enough, I will try to do better",
      "votes": null
    },
    {
      "id": "1332788",
      "postDate": "06/02/2021 09:52:28",
      "content": "<p>I am so happy you stayed near the top, I was rooting for you and other single contestant teams.  Congrats for the solo gold.  You motivated me to go higher, and your previous birdsong solution helped me as well.</p>\n<p>Your result shows that even with a single GPU one can be competitive in deep learning competitions.  This is awesome.</p>",
      "rawMarkdown": "I am so happy you stayed near the top, I was rooting for you and other single contestant teams.  Congrats for the solo gold.  You motivated me to go higher, and your previous birdsong solution helped me as well.\n\nYour result shows that even with a single GPU one can be competitive in deep learning competitions.  This is awesome.",
      "votes": null
    },
    {
      "id": "1332796",
      "postDate": "06/02/2021 09:59:43",
      "content": "<p>Thank you. I look forward to seeing you in other competitions</p>",
      "rawMarkdown": "Thank you. I look forward to seeing you in other competitions",
      "votes": null
    },
    {
      "id": "1332801",
      "postDate": "06/02/2021 10:02:16",
      "content": "<p>I will probably join  SETI, see you there ;)</p>",
      "rawMarkdown": "I will probably join  SETI, see you there ;)",
      "votes": null
    },
    {
      "id": "1333025",
      "postDate": "06/02/2021 12:44:15",
      "content": "<p>Thanks for the write up and congrats on your strong solo finish! </p>",
      "rawMarkdown": "Thanks for the write up and congrats on your strong solo finish!",
      "votes": null
    },
    {
      "id": "1333347",
      "postDate": "06/02/2021 16:41:39",
      "content": "<p>Congrats on 5th place and solo gold! :) Well done!</p>",
      "rawMarkdown": "Congrats on 5th place and solo gold! :) Well done!",
      "votes": null
    },
    {
      "id": "1333751",
      "postDate": "06/03/2021 03:21:33",
      "content": "<p>Congrats the strong finish and another solo gold <a href=\"https://www.kaggle.com/vlomme\" target=\"_blank\">@vlomme</a> , our strongest model also benefited a lot from you previous Cornell bird competition, thanks a lot. <br>\nOnly I have one question about your random_power(contrast) augmentation, when I build my model at the early stage this contract augmentation indeed improve my score a lot(+0.02 CV) on my single fold model, but when I tried to train all 5 folds model and ensemble, I found the ensemble result only improve very little, so I investigate my augmentation with visualization, I found your random_power(0.5, 3.5) have very high probability to make images very quiet(only bird call is visible, background noise almost not seen), so I doubt this is the reason that cause my 5 folds models loss some diversity thus ensemble result is not so good. I am wondering if you have the same observation, or this random_power is always good for you? </p>",
      "rawMarkdown": "Congrats the strong finish and another solo gold @vlomme , our strongest model also benefited a lot from you previous Cornell bird competition, thanks a lot. \nOnly I have one question about your random_power(contrast) augmentation, when I build my model at the early stage this contract augmentation indeed improve my score a lot(+0.02 CV) on my single fold model, but when I tried to train all 5 folds model and ensemble, I found the ensemble result only improve very little, so I investigate my augmentation with visualization, I found your random_power(0.5, 3.5) have very high probability to make images very quiet(only bird call is visible, background noise almost not seen), so I doubt this is the reason that cause my 5 folds models loss some diversity thus ensemble result is not so good. I am wondering if you have the same observation, or this random_power is always good for you?",
      "votes": null
    },
    {
      "id": "1333874",
      "postDate": "06/03/2021 05:37:33",
      "content": "<p>Thank you. For me, random_power gave an improvement as well. Probably best to mix models with and without random_power </p>",
      "rawMarkdown": "Thank you. For me, random_power gave an improvement as well. Probably best to mix models with and without random_power",
      "votes": null
    },
    {
      "id": "1333920",
      "postDate": "06/03/2021 06:34:29",
      "content": "<p>Thanks for replay, your augmentation code with spectrogram is very useful and straightforward, in my another model I tried augmentation with audiomentation library on wave, but result is not as good as yours.</p>\n<p>And I saw you did random_power() twice, one before power_to_db() and one afterward, I tried remove one of them but the score degrade, could you please explain a bit that why do random_power() twice is better? Thanks.</p>",
      "rawMarkdown": "Thanks for replay, your augmentation code with spectrogram is very useful and straightforward, in my another model I tried augmentation with audiomentation library on wave, but result is not as good as yours.\n\nAnd I saw you did random_power() twice, one before power_to_db() and one afterward, I tried remove one of them but the score degrade, could you please explain a bit that why do random_power() twice is better? Thanks.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1332788,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "06/02/2021 09:52:28",
      "content": "<p>I am so happy you stayed near the top, I was rooting for you and other single contestant teams.  Congrats for the solo gold.  You motivated me to go higher, and your previous birdsong solution helped me as well.</p>\n<p>Your result shows that even with a single GPU one can be competitive in deep learning competitions.  This is awesome.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1332796,
          "author_name": "vlomme",
          "author_url": "",
          "post_date": "06/02/2021 09:59:43",
          "content": "<p>Thank you. I look forward to seeing you in other competitions</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1332801,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "06/02/2021 10:02:16",
          "content": "<p>I will probably join  SETI, see you there ;)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1333025,
      "author_name": "sgalib",
      "author_url": "",
      "post_date": "06/02/2021 12:44:15",
      "content": "<p>Thanks for the write up and congrats on your strong solo finish! </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1333347,
      "author_name": "piantic",
      "author_url": "",
      "post_date": "06/02/2021 16:41:39",
      "content": "<p>Congrats on 5th place and solo gold! :) Well done!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1333751,
      "author_name": "superchenhao",
      "author_url": "",
      "post_date": "06/03/2021 03:21:33",
      "content": "<p>Congrats the strong finish and another solo gold <a href=\"https://www.kaggle.com/vlomme\" target=\"_blank\">@vlomme</a> , our strongest model also benefited a lot from you previous Cornell bird competition, thanks a lot. <br>\nOnly I have one question about your random_power(contrast) augmentation, when I build my model at the early stage this contract augmentation indeed improve my score a lot(+0.02 CV) on my single fold model, but when I tried to train all 5 folds model and ensemble, I found the ensemble result only improve very little, so I investigate my augmentation with visualization, I found your random_power(0.5, 3.5) have very high probability to make images very quiet(only bird call is visible, background noise almost not seen), so I doubt this is the reason that cause my 5 folds models loss some diversity thus ensemble result is not so good. I am wondering if you have the same observation, or this random_power is always good for you? </p>",
      "votes": null,
      "replies": [
        {
          "id": 1333874,
          "author_name": "vlomme",
          "author_url": "",
          "post_date": "06/03/2021 05:37:33",
          "content": "<p>Thank you. For me, random_power gave an improvement as well. Probably best to mix models with and without random_power </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1333920,
          "author_name": "superchenhao",
          "author_url": "",
          "post_date": "06/03/2021 06:34:29",
          "content": "<p>Thanks for replay, your augmentation code with spectrogram is very useful and straightforward, in my another model I tried augmentation with audiomentation library on wave, but result is not as good as yours.</p>\n<p>And I saw you did random_power() twice, one before power_to_db() and one afterward, I tried remove one of them but the score degrade, could you please explain a bit that why do random_power() twice is better? Thanks.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1332514": "Thanks to the organizers for an interesting contest and @cpmpml for the motivation.\n\n**My decision is based on my public code posted here:**\nhttps://www.kaggle.com/c/birdsong-recognition/discussion/183269\nhttps://github.com/vlomme/Birdcall-Identification-competition\n\n**The basic solution easily gets silver. What I changed:**\n- Switched to SED +1%.\n- Lowered threshold +1% . \n- Left only birds from region of record +0%\n- Changed ensemble averaging +1%\n\n## Pre-processing:\nUsed log-melspectrograms\nn_fft = 1536, sr = 21952, hop_length = 245, n_mels = 224, len_chack 448, image_size = 224 * 448, 5 seconds\n\n## Models:\nEnsemble of 14 models (sed_resnet50, sed_resnest50, sed _efficientnet-b0)\n\n## Augmentations:\n- For contrast, I raised the image to a power of 0.5 to 3. at 0.5, the background noise is closer to the birds, and at 3, on the contrary, the quiet sounds become even quieter.\n- Slightly accelerated / slowed down recording\n- Add a different sound without birds(rain, noise, conversations, etc.)\n- Added white, pink, and band noise. Increasing the noise level increases recall, but reduces precision.\n- With a probability of 0.5 lowered the upper frequencies. In the real world, the upper frequencies fade faster with distance\n\n## Train:\n- Used BCEWithLogitsLoss. For the main birds, the label was 1. For birds in the background 0.3.\n- Used loss:\n```\ntrain_los1 = nn.BCEWithLogitsLoss()(prediction['clipwise_output'], true)\ntrain_los2 = nn.BCEWithLogitsLoss()(prediction[\"segmentwise_output_max\"], true)\ntrain_loss = (train_los1 + train_los2)/2\n```\n- I didn't look at metrics on training records, but only on validation files (train_soundscapes)\n- 20-40 epochs (24 hours on gtx1060)\n\n## Postprocessing\n- If there was a bird in the segment, I increased the probability of finding it in the entire file.\n- Model ensemble averaging\n```\nproba1 = proba.prod(axis = 0) ** (1.0/len(proba))\nproba = proba**2\nproba = proba.mean(axis=0)\nproba = proba**(1/2)\nproba = (proba + proba1)/2\n```\n\nSince the training took a long time and other approaches didn't work, I switched to other competitions and have hardly taught any new models in the last month. It's a shame that a little was not enough, I will try to do better",
    "1332788": "I am so happy you stayed near the top, I was rooting for you and other single contestant teams.  Congrats for the solo gold.  You motivated me to go higher, and your previous birdsong solution helped me as well.\n\nYour result shows that even with a single GPU one can be competitive in deep learning competitions.  This is awesome.",
    "1332796": "Thank you. I look forward to seeing you in other competitions",
    "1332801": "I will probably join  SETI, see you there ;)",
    "1333025": "Thanks for the write up and congrats on your strong solo finish!",
    "1333347": "Congrats on 5th place and solo gold! :) Well done!",
    "1333751": "Congrats the strong finish and another solo gold @vlomme , our strongest model also benefited a lot from you previous Cornell bird competition, thanks a lot. \nOnly I have one question about your random_power(contrast) augmentation, when I build my model at the early stage this contract augmentation indeed improve my score a lot(+0.02 CV) on my single fold model, but when I tried to train all 5 folds model and ensemble, I found the ensemble result only improve very little, so I investigate my augmentation with visualization, I found your random_power(0.5, 3.5) have very high probability to make images very quiet(only bird call is visible, background noise almost not seen), so I doubt this is the reason that cause my 5 folds models loss some diversity thus ensemble result is not so good. I am wondering if you have the same observation, or this random_power is always good for you?",
    "1333874": "Thank you. For me, random_power gave an improvement as well. Probably best to mix models with and without random_power",
    "1333920": "Thanks for replay, your augmentation code with spectrogram is very useful and straightforward, in my another model I tried augmentation with audiomentation library on wave, but result is not as good as yours.\n\nAnd I saw you did random_power() twice, one before power_to_db() and one afterward, I tried remove one of them but the score degrade, could you please explain a bit that why do random_power() twice is better? Thanks."
  },
  "source": "meta"
}