{
  "id": 183208,
  "title": "1st Place Solution",
  "url": "/competitions/birdsong-recognition/writeups/ryan-wong-1st-place-solution",
  "author_name": "",
  "post_date": "2020-09-21T13:28:41.313Z",
  "votes": 147,
  "comment_count": 33,
  "views": 0,
  "content": "<p>Most of my solution was based on the baseline SED model provided by <a href=\"https://www.kaggle.com/hidehisaarai1213\" target=\"_blank\">@hidehisaarai1213</a> . Without his kernel I wouldn't have achieved the result I did. So I am really grateful to him. Thanks for sharing a lot during the competition, I learnt a lot. </p>\n<h2>Data Augmentation</h2>\n<p>No external data.</p>\n<ul>\n<li>Pink noise</li>\n<li>Gaussian noise</li>\n<li>Gaussian SNR</li>\n<li>Gain (Volume Adjustment)</li>\n</ul>\n<h2>Models</h2>\n<p>I noticed that the default SED model had over 80 million parameters so I switched all my models to use a pretrained densenet121 model as the cnn feature extractor and reduced the attention block size to 1024. Since it was much smaller and wouldn't overfit as much as we only had around 100 files for each audio class. I mainly tried densenet as previous top solutions to audio competitions used a densenet like architecture. I also replaced the clamp on the attention with tanh as mentioned in the <a href=\"https://www.kaggle.com/hidehisaarai1213/introduction-to-sound-event-detection/comments#962915\" target=\"_blank\">comments on the SED notebook</a></p>\n<ul>\n<li>4 fold models without mixup</li>\n<li>4 fold models with mixup</li>\n<li>5 fold models without mixup</li>\n</ul>\n<h2>Training</h2>\n<ul>\n<li>Cosine Annealing Scheduler with warmup </li>\n<li>batch size of 28</li>\n<li>Mixup (on 4 of the final models)</li>\n<li>50 epochs for non-mixup models and 100 epochs for mixup models</li>\n<li>AdamW with weight_decay 0.01</li>\n<li>SpecAugmentation enabled</li>\n<li>30 second audio clips during training and evaluating on 2 30 second clips per audio.</li>\n</ul>\n<h3>Loss Function</h3>\n<p>My loss function looked something like the below. I wanted to experiment with different parameters but in the end I mainly used the default values, which was just BCELoss. I used a different loss function for 2 of the non-mixup models and it was based on randomly removing the primary label predictions from the loss function, to try increase the secondary_label predictions but I gave up on the approach for the rest of the models since I was running out of time and resources.</p>\n<pre><code>class SedScaledPosNegFocalLoss(nn.Module):\n    def __init__(self, gamma=0.0, alpha_1=1.0, alpha_0=1.0, secondary_factor=1.0):\n        super().__init__()\n\n        self.loss_fn = nn.BCELoss(reduction='none')\n        self.secondary_factor = secondary_factor\n        self.gamma = gamma\n        self.alpha_1 = alpha_1\n        self.alpha_0 = alpha_0\n        self.loss_keys = [\"bce_loss\", \"F_loss\", \"FScaled_loss\", \"F_loss_0\", \"F_loss_1\"]\n\n    def forward(self, y_pred, y_target):\n        y_true = y_target[\"all_labels\"]\n        y_sec_true = y_target[\"secondary_labels\"]\n        bs, s, o = y_true.shape\n\n        # Sigmoid has already been applied in the model\n        y_pred = torch.clamp(y_pred, min=EPSILON_FP16, max=1.0-EPSILON_FP16)\n        y_pred = y_pred.reshape(bs*s,o)\n        y_true = y_true.reshape(bs*s,o)\n        y_sec_true = y_sec_true.reshape(bs*s,o)\n\n        with torch.no_grad():\n            y_all_ones_mask = torch.ones_like(y_true, requires_grad=False)\n            y_all_zeros_mask = torch.zeros_like(y_true, requires_grad=False)\n            y_all_mask = torch.where(y_true &gt; 0.0, y_all_ones_mask, y_all_zeros_mask)\n            y_ones_mask = torch.ones_like(y_sec_true, requires_grad=False)\n            y_zeros_mask = torch.ones_like(y_sec_true, requires_grad=False) *self.secondary_factor\n            y_secondary_mask = torch.where(y_sec_true &gt; 0.0, y_zeros_mask, y_ones_mask)\n        bce_loss = self.loss_fn(y_pred, y_true)\n        pt = torch.exp(-bce_loss)\n        F_loss_0 = (self.alpha_0*(1-y_all_mask)) * (1-pt)**self.gamma * bce_loss\n        F_loss_1 = (self.alpha_1*y_all_mask) * (1-pt)**self.gamma * bce_loss\n\n        F_loss = F_loss_0 + F_loss_1\n\n        FScaled_loss = y_secondary_mask*F_loss\n        FScaled_loss = FScaled_loss.mean()\n\n        return FScaled_loss, {\"bce_loss\": bce_loss.mean(), \"F_loss_1\": F_loss_1.mean(), \"F_loss_0\": F_loss_0.mean(), \"F_loss\": F_loss.mean(), \"FScaled_loss\": FScaled_loss }\n</code></pre>\n<p>`</p>\n<h2>Thresholds</h2>\n<p>I used a threshold of 0.3 on the <code>framewise_output</code> and 0.3 on the <code>clipwise_output</code> to reduce the impact of false positives. So if the 30 second clip contained a bird according to the clipwise prediction and the 5 second interval based on framewise prediction also said it had the same bird then it would be a valid prediction. During inference I also applied 10 TTA by just adding the same audio sample 10 times in the batch and enabling Spec Augmentation.</p>\n<h2>CV vs LB</h2>\n<p>My CV didn't match the public LB at all, so I mainly relied on the LB for feedback. During training I monitored the f1 score of the clipwise prediction, framewise prediction and the loss associated with classes existing in the audio (i.e the value of <code>F_loss_1</code> in the above loss function). When loss value of <code>F_loss_1</code> increased it generally meant that it would do worse on the LB even though the f1 score was increasing too. </p>\n<h2>Ensemble</h2>\n<p>I used voting to ensemble the models. My voting selection was based on LB score so in total I had 13 models with 4 votes to consider if the bird existed or not. <br>\nOn the public LB, the 3 votes approach scored 0.617 which was slightly better than 4 votes model of 0.616, but I didn't select the 3 votes approach as I thought it was too risky which turnout out to be the correct choice as the 4 votes approach achieved 0.002 better than the 3 votes model on the private LB. My second selected submission was an an ensemble of the nomix up models (9 models) with 3 votes which scored 0.676 private, 0.613 public LB.</p>\n<p>My individual models were pretty bad on the public LB. I didn't check some of them individually as I was running out of submissions but they generally ranged between 0.585-0.605 on the Public LB.  I mainly relied on my ensemble technique to get the score boost.</p>\n<p>Thanks to the hosts and Kaggle for this interesting competition. </p>\n<p><strong>Inference Notebook</strong>: <a href=\"https://www.kaggle.com/taggatle/cornell-birdcall-identification-1st-place-solution\" target=\"_blank\">https://www.kaggle.com/taggatle/cornell-birdcall-identification-1st-place-solution</a> <br>\n<strong>Training Code</strong>: <a href=\"https://github.com/ryanwongsa/kaggle-birdsong-recognition\" target=\"_blank\">https://github.com/ryanwongsa/kaggle-birdsong-recognition</a><br>\n<strong>Example on How to train the model on Kaggle Kernels</strong>: <a href=\"https://www.kaggle.com/taggatle/example-training-notebook\" target=\"_blank\">https://www.kaggle.com/taggatle/example-training-notebook</a></p>",
  "messages": [
    {
      "id": "1012186",
      "postDate": "09/16/2020 00:44:06",
      "content": "<p>Most of my solution was based on the baseline SED model provided by <a href=\"https://www.kaggle.com/hidehisaarai1213\" target=\"_blank\">@hidehisaarai1213</a> . Without his kernel I wouldn't have achieved the result I did. So I am really grateful to him. Thanks for sharing a lot during the competition, I learnt a lot. </p>\n<h2>Data Augmentation</h2>\n<p>No external data.</p>\n<ul>\n<li>Pink noise</li>\n<li>Gaussian noise</li>\n<li>Gaussian SNR</li>\n<li>Gain (Volume Adjustment)</li>\n</ul>\n<h2>Models</h2>\n<p>I noticed that the default SED model had over 80 million parameters so I switched all my models to use a pretrained densenet121 model as the cnn feature extractor and reduced the attention block size to 1024. Since it was much smaller and wouldn't overfit as much as we only had around 100 files for each audio class. I mainly tried densenet as previous top solutions to audio competitions used a densenet like architecture. I also replaced the clamp on the attention with tanh as mentioned in the <a href=\"https://www.kaggle.com/hidehisaarai1213/introduction-to-sound-event-detection/comments#962915\" target=\"_blank\">comments on the SED notebook</a></p>\n<ul>\n<li>4 fold models without mixup</li>\n<li>4 fold models with mixup</li>\n<li>5 fold models without mixup</li>\n</ul>\n<h2>Training</h2>\n<ul>\n<li>Cosine Annealing Scheduler with warmup </li>\n<li>batch size of 28</li>\n<li>Mixup (on 4 of the final models)</li>\n<li>50 epochs for non-mixup models and 100 epochs for mixup models</li>\n<li>AdamW with weight_decay 0.01</li>\n<li>SpecAugmentation enabled</li>\n<li>30 second audio clips during training and evaluating on 2 30 second clips per audio.</li>\n</ul>\n<h3>Loss Function</h3>\n<p>My loss function looked something like the below. I wanted to experiment with different parameters but in the end I mainly used the default values, which was just BCELoss. I used a different loss function for 2 of the non-mixup models and it was based on randomly removing the primary label predictions from the loss function, to try increase the secondary_label predictions but I gave up on the approach for the rest of the models since I was running out of time and resources.</p>\n<pre><code>class SedScaledPosNegFocalLoss(nn.Module):\n    def __init__(self, gamma=0.0, alpha_1=1.0, alpha_0=1.0, secondary_factor=1.0):\n        super().__init__()\n\n        self.loss_fn = nn.BCELoss(reduction='none')\n        self.secondary_factor = secondary_factor\n        self.gamma = gamma\n        self.alpha_1 = alpha_1\n        self.alpha_0 = alpha_0\n        self.loss_keys = [\"bce_loss\", \"F_loss\", \"FScaled_loss\", \"F_loss_0\", \"F_loss_1\"]\n\n    def forward(self, y_pred, y_target):\n        y_true = y_target[\"all_labels\"]\n        y_sec_true = y_target[\"secondary_labels\"]\n        bs, s, o = y_true.shape\n\n        # Sigmoid has already been applied in the model\n        y_pred = torch.clamp(y_pred, min=EPSILON_FP16, max=1.0-EPSILON_FP16)\n        y_pred = y_pred.reshape(bs*s,o)\n        y_true = y_true.reshape(bs*s,o)\n        y_sec_true = y_sec_true.reshape(bs*s,o)\n\n        with torch.no_grad():\n            y_all_ones_mask = torch.ones_like(y_true, requires_grad=False)\n            y_all_zeros_mask = torch.zeros_like(y_true, requires_grad=False)\n            y_all_mask = torch.where(y_true &gt; 0.0, y_all_ones_mask, y_all_zeros_mask)\n            y_ones_mask = torch.ones_like(y_sec_true, requires_grad=False)\n            y_zeros_mask = torch.ones_like(y_sec_true, requires_grad=False) *self.secondary_factor\n            y_secondary_mask = torch.where(y_sec_true &gt; 0.0, y_zeros_mask, y_ones_mask)\n        bce_loss = self.loss_fn(y_pred, y_true)\n        pt = torch.exp(-bce_loss)\n        F_loss_0 = (self.alpha_0*(1-y_all_mask)) * (1-pt)**self.gamma * bce_loss\n        F_loss_1 = (self.alpha_1*y_all_mask) * (1-pt)**self.gamma * bce_loss\n\n        F_loss = F_loss_0 + F_loss_1\n\n        FScaled_loss = y_secondary_mask*F_loss\n        FScaled_loss = FScaled_loss.mean()\n\n        return FScaled_loss, {\"bce_loss\": bce_loss.mean(), \"F_loss_1\": F_loss_1.mean(), \"F_loss_0\": F_loss_0.mean(), \"F_loss\": F_loss.mean(), \"FScaled_loss\": FScaled_loss }\n</code></pre>\n<p>`</p>\n<h2>Thresholds</h2>\n<p>I used a threshold of 0.3 on the <code>framewise_output</code> and 0.3 on the <code>clipwise_output</code> to reduce the impact of false positives. So if the 30 second clip contained a bird according to the clipwise prediction and the 5 second interval based on framewise prediction also said it had the same bird then it would be a valid prediction. During inference I also applied 10 TTA by just adding the same audio sample 10 times in the batch and enabling Spec Augmentation.</p>\n<h2>CV vs LB</h2>\n<p>My CV didn't match the public LB at all, so I mainly relied on the LB for feedback. During training I monitored the f1 score of the clipwise prediction, framewise prediction and the loss associated with classes existing in the audio (i.e the value of <code>F_loss_1</code> in the above loss function). When loss value of <code>F_loss_1</code> increased it generally meant that it would do worse on the LB even though the f1 score was increasing too. </p>\n<h2>Ensemble</h2>\n<p>I used voting to ensemble the models. My voting selection was based on LB score so in total I had 13 models with 4 votes to consider if the bird existed or not. <br>\nOn the public LB, the 3 votes approach scored 0.617 which was slightly better than 4 votes model of 0.616, but I didn't select the 3 votes approach as I thought it was too risky which turnout out to be the correct choice as the 4 votes approach achieved 0.002 better than the 3 votes model on the private LB. My second selected submission was an an ensemble of the nomix up models (9 models) with 3 votes which scored 0.676 private, 0.613 public LB.</p>\n<p>My individual models were pretty bad on the public LB. I didn't check some of them individually as I was running out of submissions but they generally ranged between 0.585-0.605 on the Public LB.  I mainly relied on my ensemble technique to get the score boost.</p>\n<p>Thanks to the hosts and Kaggle for this interesting competition. </p>\n<p><strong>Inference Notebook</strong>: <a href=\"https://www.kaggle.com/taggatle/cornell-birdcall-identification-1st-place-solution\" target=\"_blank\">https://www.kaggle.com/taggatle/cornell-birdcall-identification-1st-place-solution</a> <br>\n<strong>Training Code</strong>: <a href=\"https://github.com/ryanwongsa/kaggle-birdsong-recognition\" target=\"_blank\">https://github.com/ryanwongsa/kaggle-birdsong-recognition</a><br>\n<strong>Example on How to train the model on Kaggle Kernels</strong>: <a href=\"https://www.kaggle.com/taggatle/example-training-notebook\" target=\"_blank\">https://www.kaggle.com/taggatle/example-training-notebook</a></p>",
      "rawMarkdown": "Most of my solution was based on the baseline SED model provided by @hidehisaarai1213 . Without his kernel I wouldn't have achieved the result I did. So I am really grateful to him. Thanks for sharing a lot during the competition, I learnt a lot. \n\n## Data Augmentation\n\n\nNo external data.\n\n- Pink noise\n- Gaussian noise\n- Gaussian SNR\n- Gain (Volume Adjustment)\n\n## Models\n\nI noticed that the default SED model had over 80 million parameters so I switched all my models to use a pretrained densenet121 model as the cnn feature extractor and reduced the attention block size to 1024. Since it was much smaller and wouldn't overfit as much as we only had around 100 files for each audio class. I mainly tried densenet as previous top solutions to audio competitions used a densenet like architecture. I also replaced the clamp on the attention with tanh as mentioned in the [comments on the SED notebook](https://www.kaggle.com/hidehisaarai1213/introduction-to-sound-event-detection/comments#962915)\n\n- 4 fold models without mixup\n- 4 fold models with mixup\n- 5 fold models without mixup\n\n## Training\n\n- Cosine Annealing Scheduler with warmup \n- batch size of 28\n- Mixup (on 4 of the final models)\n- 50 epochs for non-mixup models and 100 epochs for mixup models\n- AdamW with weight_decay 0.01\n- SpecAugmentation enabled\n- 30 second audio clips during training and evaluating on 2 30 second clips per audio.\n\n### Loss Function\n\nMy loss function looked something like the below. I wanted to experiment with different parameters but in the end I mainly used the default values, which was just BCELoss. I used a different loss function for 2 of the non-mixup models and it was based on randomly removing the primary label predictions from the loss function, to try increase the secondary_label predictions but I gave up on the approach for the rest of the models since I was running out of time and resources.\n\n```python\nclass SedScaledPosNegFocalLoss(nn.Module):\n    def __init__(self, gamma=0.0, alpha_1=1.0, alpha_0=1.0, secondary_factor=1.0):\n        super().__init__()\n\n        self.loss_fn = nn.BCELoss(reduction='none')\n        self.secondary_factor = secondary_factor\n        self.gamma = gamma\n        self.alpha_1 = alpha_1\n        self.alpha_0 = alpha_0\n        self.loss_keys = [\"bce_loss\", \"F_loss\", \"FScaled_loss\", \"F_loss_0\", \"F_loss_1\"]\n\n    def forward(self, y_pred, y_target):\n        y_true = y_target[\"all_labels\"]\n        y_sec_true = y_target[\"secondary_labels\"]\n        bs, s, o = y_true.shape\n\n        # Sigmoid has already been applied in the model\n        y_pred = torch.clamp(y_pred, min=EPSILON_FP16, max=1.0-EPSILON_FP16)\n        y_pred = y_pred.reshape(bs*s,o)\n        y_true = y_true.reshape(bs*s,o)\n        y_sec_true = y_sec_true.reshape(bs*s,o)\n        \n        with torch.no_grad():\n            y_all_ones_mask = torch.ones_like(y_true, requires_grad=False)\n            y_all_zeros_mask = torch.zeros_like(y_true, requires_grad=False)\n            y_all_mask = torch.where(y_true > 0.0, y_all_ones_mask, y_all_zeros_mask)\n            y_ones_mask = torch.ones_like(y_sec_true, requires_grad=False)\n            y_zeros_mask = torch.ones_like(y_sec_true, requires_grad=False) *self.secondary_factor\n            y_secondary_mask = torch.where(y_sec_true > 0.0, y_zeros_mask, y_ones_mask)\n        bce_loss = self.loss_fn(y_pred, y_true)\n        pt = torch.exp(-bce_loss)\n        F_loss_0 = (self.alpha_0*(1-y_all_mask)) * (1-pt)**self.gamma * bce_loss\n        F_loss_1 = (self.alpha_1*y_all_mask) * (1-pt)**self.gamma * bce_loss\n\n        F_loss = F_loss_0 + F_loss_1\n\n        FScaled_loss = y_secondary_mask*F_loss\n        FScaled_loss = FScaled_loss.mean()\n\n        return FScaled_loss, {\"bce_loss\": bce_loss.mean(), \"F_loss_1\": F_loss_1.mean(), \"F_loss_0\": F_loss_0.mean(), \"F_loss\": F_loss.mean(), \"FScaled_loss\": FScaled_loss }\n````\n\n## Thresholds\n\nI used a threshold of 0.3 on the `framewise_output` and 0.3 on the `clipwise_output` to reduce the impact of false positives. So if the 30 second clip contained a bird according to the clipwise prediction and the 5 second interval based on framewise prediction also said it had the same bird then it would be a valid prediction. During inference I also applied 10 TTA by just adding the same audio sample 10 times in the batch and enabling Spec Augmentation.\n\n## CV vs LB \n\nMy CV didn't match the public LB at all, so I mainly relied on the LB for feedback. During training I monitored the f1 score of the clipwise prediction, framewise prediction and the loss associated with classes existing in the audio (i.e the value of `F_loss_1` in the above loss function). When loss value of `F_loss_1` increased it generally meant that it would do worse on the LB even though the f1 score was increasing too. \n\n## Ensemble\n\nI used voting to ensemble the models. My voting selection was based on LB score so in total I had 13 models with 4 votes to consider if the bird existed or not. \nOn the public LB, the 3 votes approach scored 0.617 which was slightly better than 4 votes model of 0.616, but I didn't select the 3 votes approach as I thought it was too risky which turnout out to be the correct choice as the 4 votes approach achieved 0.002 better than the 3 votes model on the private LB. My second selected submission was an an ensemble of the nomix up models (9 models) with 3 votes which scored 0.676 private, 0.613 public LB.\n\nMy individual models were pretty bad on the public LB. I didn't check some of them individually as I was running out of submissions but they generally ranged between 0.585-0.605 on the Public LB.  I mainly relied on my ensemble technique to get the score boost.\n\nThanks to the hosts and Kaggle for this interesting competition. \n\n**Inference Notebook**: https://www.kaggle.com/taggatle/cornell-birdcall-identification-1st-place-solution \n**Training Code**: https://github.com/ryanwongsa/kaggle-birdsong-recognition\n**Example on How to train the model on Kaggle Kernels**: https://www.kaggle.com/taggatle/example-training-notebook",
      "votes": null
    },
    {
      "id": "1012188",
      "postDate": "09/16/2020 00:46:26",
      "content": "<p>Congrats, Ryan! Much appreciate your participation and more information on the winning solution. </p>",
      "rawMarkdown": "Congrats, Ryan! Much appreciate your participation and more information on the winning solution.",
      "votes": null
    },
    {
      "id": "1012204",
      "postDate": "09/16/2020 00:57:56",
      "content": "<p>Happy to know SED model won this competition! It's quite interesting that you didn't use so much data augmentation.</p>\n<p>Congratz <a href=\"https://www.kaggle.com/taggatle\" target=\"_blank\">@taggatle</a> 🎉</p>",
      "rawMarkdown": "Happy to know SED model won this competition! It's quite interesting that you didn't use so much data augmentation.\n\nCongratz @taggatle 🎉",
      "votes": null
    },
    {
      "id": "1012248",
      "postDate": "09/16/2020 01:52:39",
      "content": "<p>Congrats on the solo win!  I like your voting ensemble, we could not find a way to ensemble our models.</p>",
      "rawMarkdown": "Congrats on the solo win!  I like your voting ensemble, we could not find a way to ensemble our models.",
      "votes": null
    },
    {
      "id": "1012406",
      "postDate": "09/16/2020 04:35:19",
      "content": "<p>simple and effective! Congrac)</p>",
      "rawMarkdown": "simple and effective! Congrac)",
      "votes": null
    },
    {
      "id": "1012546",
      "postDate": "09/16/2020 06:39:11",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/taggatle\" target=\"_blank\">@taggatle</a>! 🎉🎉</p>",
      "rawMarkdown": "Congrats @taggatle! 🎉🎉",
      "votes": null
    },
    {
      "id": "1012568",
      "postDate": "09/16/2020 06:56:01",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/hidehisaarai1213\" target=\"_blank\">@hidehisaarai1213</a> , I think you provided a lot of guidance and help with what you shared.</p>\n<p>I was a kind of afraid of adding too much augmentation as I didn't want to distort the bird calls too much. I mainly relied on the pink noise as suggested by one of the hosts.</p>",
      "rawMarkdown": "Thanks @hidehisaarai1213 , I think you provided a lot of guidance and help with what you shared.\n\nI was a kind of afraid of adding too much augmentation as I didn't want to distort the bird calls too much. I mainly relied on the pink noise as suggested by one of the hosts.",
      "votes": null
    },
    {
      "id": "1012569",
      "postDate": "09/16/2020 06:56:44",
      "content": "<p>Thanks, <a href=\"https://www.kaggle.com/holgerklinck\" target=\"_blank\">@holgerklinck</a>. </p>",
      "rawMarkdown": "Thanks, @holgerklinck.",
      "votes": null
    },
    {
      "id": "1012575",
      "postDate": "09/16/2020 07:00:06",
      "content": "<p>Thanks, <a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a>. Yeah I mainly used voting to reduce the false positives, I think it was the major factor in my solution since individually the models were in the silver region.</p>",
      "rawMarkdown": "Thanks, @cpmpml. Yeah I mainly used voting to reduce the false positives, I think it was the major factor in my solution since individually the models were in the silver region.",
      "votes": null
    },
    {
      "id": "1012577",
      "postDate": "09/16/2020 07:00:41",
      "content": "<p>Thanks, <a href=\"https://www.kaggle.com/kupchanski\" target=\"_blank\">@kupchanski</a></p>",
      "rawMarkdown": "Thanks, @kupchanski",
      "votes": null
    },
    {
      "id": "1012579",
      "postDate": "09/16/2020 07:00:58",
      "content": "<p>Thanks, <a href=\"https://www.kaggle.com/ramarlina\" target=\"_blank\">@ramarlina</a></p>",
      "rawMarkdown": "Thanks, @ramarlina",
      "votes": null
    },
    {
      "id": "1012688",
      "postDate": "09/16/2020 08:15:29",
      "content": "<p>Congats <a href=\"https://www.kaggle.com/taggatle\" target=\"_blank\">@taggatle</a> </p>",
      "rawMarkdown": "Congats @taggatle",
      "votes": null
    },
    {
      "id": "1012694",
      "postDate": "09/16/2020 08:18:05",
      "content": "<p>Congratulation on the win ! </p>",
      "rawMarkdown": "Congratulation on the win !",
      "votes": null
    },
    {
      "id": "1012730",
      "postDate": "09/16/2020 08:54:06",
      "content": "<p>I noticed that score of blend was the blend of scores, but I didn't conclude that averaging fold predictions wasn't great.  I will certainly try voting next time!</p>",
      "rawMarkdown": "I noticed that score of blend was the blend of scores, but I didn't conclude that averaging fold predictions wasn't great.  I will certainly try voting next time!",
      "votes": null
    },
    {
      "id": "1012906",
      "postDate": "09/16/2020 11:34:25",
      "content": "<p>Congratulations! This will be my learning material in this weekend!</p>",
      "rawMarkdown": "Congratulations! This will be my learning material in this weekend!",
      "votes": null
    },
    {
      "id": "1014919",
      "postDate": "09/17/2020 19:37:48",
      "content": "<p>Congratulation on the win !</p>",
      "rawMarkdown": "Congratulation on the win !",
      "votes": null
    },
    {
      "id": "1018815",
      "postDate": "09/20/2020 02:32:31",
      "content": "<p>Congratulations to  the 1st place <a href=\"https://www.kaggle.com/taggatle\" target=\"_blank\">@taggatle</a>!🎉</p>",
      "rawMarkdown": "Congratulations to  the 1st place @taggatle!🎉",
      "votes": null
    },
    {
      "id": "1019310",
      "postDate": "09/20/2020 10:48:53",
      "content": "<p>congratulations Ryan</p>",
      "rawMarkdown": "congratulations Ryan",
      "votes": null
    },
    {
      "id": "1020099",
      "postDate": "09/20/2020 23:04:26",
      "content": "<p>Thank you for sharing this.</p>",
      "rawMarkdown": "Thank you for sharing this.",
      "votes": null
    },
    {
      "id": "1020231",
      "postDate": "09/21/2020 03:41:22",
      "content": "<p>interesting and congratulations</p>",
      "rawMarkdown": "interesting and congratulations",
      "votes": null
    },
    {
      "id": "1020683",
      "postDate": "09/21/2020 10:56:52",
      "content": "<p>Congratulations!!!</p>",
      "rawMarkdown": "Congratulations!!!",
      "votes": null
    },
    {
      "id": "1037838",
      "postDate": "10/05/2020 11:25:49",
      "content": "<p>Could you tell me what were the scores of single models you used for ensembling? I want to know how much did hard voting contribute to your score.</p>",
      "rawMarkdown": "Could you tell me what were the scores of single models you used for ensembling? I want to know how much did hard voting contribute to your score.",
      "votes": null
    },
    {
      "id": "1038194",
      "postDate": "10/05/2020 16:14:33",
      "content": "<p>I don't have all of them as I was limited on submissions but I have an example of 3 of the results on fold 0,1,2 in the attached image. At the top contains an ensemble of 5 of the folds with 3 votes. I didn't trust the LB completely so didn't want to pick individual models based on the LB score and instead based it on the ensemble score. </p>\n<p>An ensemble of 9 models with 3 votes got private LB 0.676 and public LB 0.613, so just adding 4 more models with the same vote count scored a bit better on the private and public LB. An ensemble of 13 models with the same votes also improves the private LB to 0.679. </p>\n<p>It is possible to get a high scoring single model though. One of my earlier models when I was hyperparameter testing got a score of 0.605 on the public LB and 0.674 on the private LB. It wasn't included in the final ensemble.</p>",
      "rawMarkdown": "I don't have all of them as I was limited on submissions but I have an example of 3 of the results on fold 0,1,2 in the attached image. At the top contains an ensemble of 5 of the folds with 3 votes. I didn't trust the LB completely so didn't want to pick individual models based on the LB score and instead based it on the ensemble score. \n\nAn ensemble of 9 models with 3 votes got private LB 0.676 and public LB 0.613, so just adding 4 more models with the same vote count scored a bit better on the private and public LB. An ensemble of 13 models with the same votes also improves the private LB to 0.679. \n\nIt is possible to get a high scoring single model though. One of my earlier models when I was hyperparameter testing got a score of 0.605 on the public LB and 0.674 on the private LB. It wasn't included in the final ensemble.",
      "votes": null
    },
    {
      "id": "1043202",
      "postDate": "10/08/2020 18:57:34",
      "content": "<p>Congrats on the win!<br>\nThanks for the clear write-up :)</p>",
      "rawMarkdown": "Congrats on the win!\nThanks for the clear write-up :)",
      "votes": null
    },
    {
      "id": "1043802",
      "postDate": "10/09/2020 08:51:47",
      "content": "<p>Congratulation on the win !why you switch your models?   We could not find a way to use.</p>",
      "rawMarkdown": "Congratulation on the win !why you switch your models?   We could not find a way to use.",
      "votes": null
    },
    {
      "id": "1044666",
      "postDate": "10/10/2020 02:56:22",
      "content": "<p>good one using the pre-trained model</p>",
      "rawMarkdown": "good one using the pre-trained model",
      "votes": null
    },
    {
      "id": "1047056",
      "postDate": "10/12/2020 08:10:39",
      "content": "<p>Sorry for not replying…<br>\nSeems it's private score (also public score) varies a lot - maybe it helped the result to be good in terms of diversity.<br>\nThanks for telling that!</p>",
      "rawMarkdown": "Sorry for not replying...\nSeems it's private score (also public score) varies a lot - maybe it helped the result to be good in terms of diversity.\nThanks for telling that!",
      "votes": null
    },
    {
      "id": "1086453",
      "postDate": "11/21/2020 17:39:23",
      "content": "<p>Hi there real appreciate you sharing all you work, I was wondering if you can help me.</p>\n<p>trying to run 'train.py' but every time I get to : parser.add_argument('--config', type=str)</p>\n<p>I get this error, as the arguments are None…</p>\n<p>_StoreAction(option_strings=['--config'], dest='config', nargs=None, const=None, default=None, type=, choices=None, help=None, metavar=None)</p>\n<p>Thanks!</p>",
      "rawMarkdown": "Hi there real appreciate you sharing all you work, I was wondering if you can help me.\n\ntrying to run 'train.py' but every time I get to : parser.add_argument('--config', type=str)\n\nI get this error, as the arguments are None...\n\n _StoreAction(option_strings=['--config'], dest='config', nargs=None, const=None, default=None, type=<class 'str'>, choices=None, help=None, metavar=None)\n\nThanks!",
      "votes": null
    },
    {
      "id": "1086528",
      "postDate": "11/21/2020 18:49:17",
      "content": "<p>Hi, if you create a file for hyperparameters like the one in the Kaggle notebook (<a href=\"https://www.kaggle.com/taggatle/example-training-notebook\" target=\"_blank\">https://www.kaggle.com/taggatle/example-training-notebook</a>) and save it somewhere as \"config_params/example_config.py\". Then you can use those configs for training with the following command:</p>\n<pre><code>python sed_train.py --config \"config_params.example_config\"\n</code></pre>",
      "rawMarkdown": "Hi, if you create a file for hyperparameters like the one in the Kaggle notebook (https://www.kaggle.com/taggatle/example-training-notebook) and save it somewhere as \"config_params/example_config.py\". Then you can use those configs for training with the following command:\n```\npython sed_train.py --config \"config_params.example_config\"\n```",
      "votes": null
    },
    {
      "id": "1086539",
      "postDate": "11/21/2020 18:55:45",
      "content": "<p>Hi thanks for your replay!! really appreciate it.</p>\n<p>I was trying this but this not worked…</p>\n<p>parser.add_argument('--config', default='./config_params/configs.py', type=str)</p>\n<p>Now I see how, thanks!!</p>",
      "rawMarkdown": "Hi thanks for your replay!! really appreciate it.\n\nI was trying this but this not worked...\n\nparser.add_argument('--config', default='./config_params/configs.py', type=str)\n\nNow I see how, thanks!!",
      "votes": null
    },
    {
      "id": "1142667",
      "postDate": "01/07/2021 14:29:36",
      "content": "<p>Thanks so much for sharing your great work! Wondering, why do you think you and other teams did not make use of the learnable wavegram component of the PANN paper?</p>",
      "rawMarkdown": "Thanks so much for sharing your great work! Wondering, why do you think you and other teams did not make use of the learnable wavegram component of the PANN paper?",
      "votes": null
    },
    {
      "id": "1247882",
      "postDate": "03/22/2021 06:42:37",
      "content": "<p><a href=\"https://www.kaggle.com/taggatle\" target=\"_blank\">@taggatle</a> very late here but any chance you could explain the intuition behind your custom loss function?</p>",
      "rawMarkdown": "taggatle very late here but any chance you could explain the intuition behind your custom loss function?",
      "votes": null
    },
    {
      "id": "1254584",
      "postDate": "03/27/2021 20:41:48",
      "content": "<p>Thanks for sharing as we use this to identify birds in our backyard.</p>",
      "rawMarkdown": "Thanks for sharing as we use this to identify birds in our backyard.",
      "votes": null
    },
    {
      "id": "1306096",
      "postDate": "05/13/2021 15:58:52",
      "content": "<p><a href=\"https://www.kaggle.com/taggatle\" target=\"_blank\">@taggatle</a> will spectogram look different to model (hence difference result) if train  on more duration clips chunk and infer on less duration clip chunk say 5 sec clip chunk</p>",
      "rawMarkdown": "taggatle will spectogram look different to model (hence difference result) if train  on more duration clips chunk and infer on less duration clip chunk say 5 sec clip chunk",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1012188,
      "author_name": "holgerklinck",
      "author_url": "",
      "post_date": "09/16/2020 00:46:26",
      "content": "<p>Congrats, Ryan! Much appreciate your participation and more information on the winning solution. </p>",
      "votes": null,
      "replies": [
        {
          "id": 1012569,
          "author_name": "taggatle",
          "author_url": "",
          "post_date": "09/16/2020 06:56:44",
          "content": "<p>Thanks, <a href=\"https://www.kaggle.com/holgerklinck\" target=\"_blank\">@holgerklinck</a>. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1012204,
      "author_name": "hidehisaarai1213",
      "author_url": "",
      "post_date": "09/16/2020 00:57:56",
      "content": "<p>Happy to know SED model won this competition! It's quite interesting that you didn't use so much data augmentation.</p>\n<p>Congratz <a href=\"https://www.kaggle.com/taggatle\" target=\"_blank\">@taggatle</a> 🎉</p>",
      "votes": null,
      "replies": [
        {
          "id": 1012568,
          "author_name": "taggatle",
          "author_url": "",
          "post_date": "09/16/2020 06:56:01",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/hidehisaarai1213\" target=\"_blank\">@hidehisaarai1213</a> , I think you provided a lot of guidance and help with what you shared.</p>\n<p>I was a kind of afraid of adding too much augmentation as I didn't want to distort the bird calls too much. I mainly relied on the pink noise as suggested by one of the hosts.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1037838,
          "author_name": "hidehisaarai1213",
          "author_url": "",
          "post_date": "10/05/2020 11:25:49",
          "content": "<p>Could you tell me what were the scores of single models you used for ensembling? I want to know how much did hard voting contribute to your score.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1038194,
          "author_name": "taggatle",
          "author_url": "",
          "post_date": "10/05/2020 16:14:33",
          "content": "<p>I don't have all of them as I was limited on submissions but I have an example of 3 of the results on fold 0,1,2 in the attached image. At the top contains an ensemble of 5 of the folds with 3 votes. I didn't trust the LB completely so didn't want to pick individual models based on the LB score and instead based it on the ensemble score. </p>\n<p>An ensemble of 9 models with 3 votes got private LB 0.676 and public LB 0.613, so just adding 4 more models with the same vote count scored a bit better on the private and public LB. An ensemble of 13 models with the same votes also improves the private LB to 0.679. </p>\n<p>It is possible to get a high scoring single model though. One of my earlier models when I was hyperparameter testing got a score of 0.605 on the public LB and 0.674 on the private LB. It wasn't included in the final ensemble.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1047056,
          "author_name": "hidehisaarai1213",
          "author_url": "",
          "post_date": "10/12/2020 08:10:39",
          "content": "<p>Sorry for not replying…<br>\nSeems it's private score (also public score) varies a lot - maybe it helped the result to be good in terms of diversity.<br>\nThanks for telling that!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1012248,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "09/16/2020 01:52:39",
      "content": "<p>Congrats on the solo win!  I like your voting ensemble, we could not find a way to ensemble our models.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1012575,
          "author_name": "taggatle",
          "author_url": "",
          "post_date": "09/16/2020 07:00:06",
          "content": "<p>Thanks, <a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a>. Yeah I mainly used voting to reduce the false positives, I think it was the major factor in my solution since individually the models were in the silver region.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1012730,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "09/16/2020 08:54:06",
          "content": "<p>I noticed that score of blend was the blend of scores, but I didn't conclude that averaging fold predictions wasn't great.  I will certainly try voting next time!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1012406,
      "author_name": "kupchanski",
      "author_url": "",
      "post_date": "09/16/2020 04:35:19",
      "content": "<p>simple and effective! Congrac)</p>",
      "votes": null,
      "replies": [
        {
          "id": 1012577,
          "author_name": "taggatle",
          "author_url": "",
          "post_date": "09/16/2020 07:00:41",
          "content": "<p>Thanks, <a href=\"https://www.kaggle.com/kupchanski\" target=\"_blank\">@kupchanski</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1012546,
      "author_name": "ramarlina",
      "author_url": "",
      "post_date": "09/16/2020 06:39:11",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/taggatle\" target=\"_blank\">@taggatle</a>! 🎉🎉</p>",
      "votes": null,
      "replies": [
        {
          "id": 1012579,
          "author_name": "taggatle",
          "author_url": "",
          "post_date": "09/16/2020 07:00:58",
          "content": "<p>Thanks, <a href=\"https://www.kaggle.com/ramarlina\" target=\"_blank\">@ramarlina</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1012688,
      "author_name": "doanquanvietnamca",
      "author_url": "",
      "post_date": "09/16/2020 08:15:29",
      "content": "<p>Congats <a href=\"https://www.kaggle.com/taggatle\" target=\"_blank\">@taggatle</a> </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1012694,
      "author_name": "theoviel",
      "author_url": "",
      "post_date": "09/16/2020 08:18:05",
      "content": "<p>Congratulation on the win ! </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1012906,
      "author_name": "fiyeroleung",
      "author_url": "",
      "post_date": "09/16/2020 11:34:25",
      "content": "<p>Congratulations! This will be my learning material in this weekend!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1014919,
      "author_name": "marek3000",
      "author_url": "",
      "post_date": "09/17/2020 19:37:48",
      "content": "<p>Congratulation on the win !</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1018815,
      "author_name": "hsinwenchang",
      "author_url": "",
      "post_date": "09/20/2020 02:32:31",
      "content": "<p>Congratulations to  the 1st place <a href=\"https://www.kaggle.com/taggatle\" target=\"_blank\">@taggatle</a>!🎉</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1020099,
      "author_name": "zhjckfd",
      "author_url": "",
      "post_date": "09/20/2020 23:04:26",
      "content": "<p>Thank you for sharing this.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1020683,
      "author_name": "tyadav",
      "author_url": "",
      "post_date": "09/21/2020 10:56:52",
      "content": "<p>Congratulations!!!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1043202,
      "author_name": "xenophule",
      "author_url": "",
      "post_date": "10/08/2020 18:57:34",
      "content": "<p>Congrats on the win!<br>\nThanks for the clear write-up :)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1043802,
      "author_name": "minghsiuyo",
      "author_url": "",
      "post_date": "10/09/2020 08:51:47",
      "content": "<p>Congratulation on the win !why you switch your models?   We could not find a way to use.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1044666,
      "author_name": "aaquilamariajohn",
      "author_url": "",
      "post_date": "10/10/2020 02:56:22",
      "content": "<p>good one using the pre-trained model</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1086453,
      "author_name": "oscarrangel",
      "author_url": "",
      "post_date": "11/21/2020 17:39:23",
      "content": "<p>Hi there real appreciate you sharing all you work, I was wondering if you can help me.</p>\n<p>trying to run 'train.py' but every time I get to : parser.add_argument('--config', type=str)</p>\n<p>I get this error, as the arguments are None…</p>\n<p>_StoreAction(option_strings=['--config'], dest='config', nargs=None, const=None, default=None, type=, choices=None, help=None, metavar=None)</p>\n<p>Thanks!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1086528,
          "author_name": "taggatle",
          "author_url": "",
          "post_date": "11/21/2020 18:49:17",
          "content": "<p>Hi, if you create a file for hyperparameters like the one in the Kaggle notebook (<a href=\"https://www.kaggle.com/taggatle/example-training-notebook\" target=\"_blank\">https://www.kaggle.com/taggatle/example-training-notebook</a>) and save it somewhere as \"config_params/example_config.py\". Then you can use those configs for training with the following command:</p>\n<pre><code>python sed_train.py --config \"config_params.example_config\"\n</code></pre>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1086539,
          "author_name": "oscarrangel",
          "author_url": "",
          "post_date": "11/21/2020 18:55:45",
          "content": "<p>Hi thanks for your replay!! really appreciate it.</p>\n<p>I was trying this but this not worked…</p>\n<p>parser.add_argument('--config', default='./config_params/configs.py', type=str)</p>\n<p>Now I see how, thanks!!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1142667,
      "author_name": "alexandersoare",
      "author_url": "",
      "post_date": "01/07/2021 14:29:36",
      "content": "<p>Thanks so much for sharing your great work! Wondering, why do you think you and other teams did not make use of the learnable wavegram component of the PANN paper?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1247882,
      "author_name": "aryaman1999",
      "author_url": "",
      "post_date": "03/22/2021 06:42:37",
      "content": "<p><a href=\"https://www.kaggle.com/taggatle\" target=\"_blank\">@taggatle</a> very late here but any chance you could explain the intuition behind your custom loss function?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1254584,
      "author_name": "tomkkk",
      "author_url": "",
      "post_date": "03/27/2021 20:41:48",
      "content": "<p>Thanks for sharing as we use this to identify birds in our backyard.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1306096,
      "author_name": "jaideepvalani",
      "author_url": "",
      "post_date": "05/13/2021 15:58:52",
      "content": "<p><a href=\"https://www.kaggle.com/taggatle\" target=\"_blank\">@taggatle</a> will spectogram look different to model (hence difference result) if train  on more duration clips chunk and infer on less duration clip chunk say 5 sec clip chunk</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1019310,
      "author_name": "lokeswarreddy",
      "author_url": "",
      "post_date": "09/20/2020 10:48:53",
      "content": "<p>congratulations Ryan</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1020231,
      "author_name": "aamirjahan",
      "author_url": "",
      "post_date": "09/21/2020 03:41:22",
      "content": "<p>interesting and congratulations</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1012186": "Most of my solution was based on the baseline SED model provided by @hidehisaarai1213 . Without his kernel I wouldn't have achieved the result I did. So I am really grateful to him. Thanks for sharing a lot during the competition, I learnt a lot. \n\n## Data Augmentation\n\n\nNo external data.\n\n- Pink noise\n- Gaussian noise\n- Gaussian SNR\n- Gain (Volume Adjustment)\n\n## Models\n\nI noticed that the default SED model had over 80 million parameters so I switched all my models to use a pretrained densenet121 model as the cnn feature extractor and reduced the attention block size to 1024. Since it was much smaller and wouldn't overfit as much as we only had around 100 files for each audio class. I mainly tried densenet as previous top solutions to audio competitions used a densenet like architecture. I also replaced the clamp on the attention with tanh as mentioned in the [comments on the SED notebook](https://www.kaggle.com/hidehisaarai1213/introduction-to-sound-event-detection/comments#962915)\n\n- 4 fold models without mixup\n- 4 fold models with mixup\n- 5 fold models without mixup\n\n## Training\n\n- Cosine Annealing Scheduler with warmup \n- batch size of 28\n- Mixup (on 4 of the final models)\n- 50 epochs for non-mixup models and 100 epochs for mixup models\n- AdamW with weight_decay 0.01\n- SpecAugmentation enabled\n- 30 second audio clips during training and evaluating on 2 30 second clips per audio.\n\n### Loss Function\n\nMy loss function looked something like the below. I wanted to experiment with different parameters but in the end I mainly used the default values, which was just BCELoss. I used a different loss function for 2 of the non-mixup models and it was based on randomly removing the primary label predictions from the loss function, to try increase the secondary_label predictions but I gave up on the approach for the rest of the models since I was running out of time and resources.\n\n```python\nclass SedScaledPosNegFocalLoss(nn.Module):\n    def __init__(self, gamma=0.0, alpha_1=1.0, alpha_0=1.0, secondary_factor=1.0):\n        super().__init__()\n\n        self.loss_fn = nn.BCELoss(reduction='none')\n        self.secondary_factor = secondary_factor\n        self.gamma = gamma\n        self.alpha_1 = alpha_1\n        self.alpha_0 = alpha_0\n        self.loss_keys = [\"bce_loss\", \"F_loss\", \"FScaled_loss\", \"F_loss_0\", \"F_loss_1\"]\n\n    def forward(self, y_pred, y_target):\n        y_true = y_target[\"all_labels\"]\n        y_sec_true = y_target[\"secondary_labels\"]\n        bs, s, o = y_true.shape\n\n        # Sigmoid has already been applied in the model\n        y_pred = torch.clamp(y_pred, min=EPSILON_FP16, max=1.0-EPSILON_FP16)\n        y_pred = y_pred.reshape(bs*s,o)\n        y_true = y_true.reshape(bs*s,o)\n        y_sec_true = y_sec_true.reshape(bs*s,o)\n        \n        with torch.no_grad():\n            y_all_ones_mask = torch.ones_like(y_true, requires_grad=False)\n            y_all_zeros_mask = torch.zeros_like(y_true, requires_grad=False)\n            y_all_mask = torch.where(y_true > 0.0, y_all_ones_mask, y_all_zeros_mask)\n            y_ones_mask = torch.ones_like(y_sec_true, requires_grad=False)\n            y_zeros_mask = torch.ones_like(y_sec_true, requires_grad=False) *self.secondary_factor\n            y_secondary_mask = torch.where(y_sec_true > 0.0, y_zeros_mask, y_ones_mask)\n        bce_loss = self.loss_fn(y_pred, y_true)\n        pt = torch.exp(-bce_loss)\n        F_loss_0 = (self.alpha_0*(1-y_all_mask)) * (1-pt)**self.gamma * bce_loss\n        F_loss_1 = (self.alpha_1*y_all_mask) * (1-pt)**self.gamma * bce_loss\n\n        F_loss = F_loss_0 + F_loss_1\n\n        FScaled_loss = y_secondary_mask*F_loss\n        FScaled_loss = FScaled_loss.mean()\n\n        return FScaled_loss, {\"bce_loss\": bce_loss.mean(), \"F_loss_1\": F_loss_1.mean(), \"F_loss_0\": F_loss_0.mean(), \"F_loss\": F_loss.mean(), \"FScaled_loss\": FScaled_loss }\n````\n\n## Thresholds\n\nI used a threshold of 0.3 on the `framewise_output` and 0.3 on the `clipwise_output` to reduce the impact of false positives. So if the 30 second clip contained a bird according to the clipwise prediction and the 5 second interval based on framewise prediction also said it had the same bird then it would be a valid prediction. During inference I also applied 10 TTA by just adding the same audio sample 10 times in the batch and enabling Spec Augmentation.\n\n## CV vs LB \n\nMy CV didn't match the public LB at all, so I mainly relied on the LB for feedback. During training I monitored the f1 score of the clipwise prediction, framewise prediction and the loss associated with classes existing in the audio (i.e the value of `F_loss_1` in the above loss function). When loss value of `F_loss_1` increased it generally meant that it would do worse on the LB even though the f1 score was increasing too. \n\n## Ensemble\n\nI used voting to ensemble the models. My voting selection was based on LB score so in total I had 13 models with 4 votes to consider if the bird existed or not. \nOn the public LB, the 3 votes approach scored 0.617 which was slightly better than 4 votes model of 0.616, but I didn't select the 3 votes approach as I thought it was too risky which turnout out to be the correct choice as the 4 votes approach achieved 0.002 better than the 3 votes model on the private LB. My second selected submission was an an ensemble of the nomix up models (9 models) with 3 votes which scored 0.676 private, 0.613 public LB.\n\nMy individual models were pretty bad on the public LB. I didn't check some of them individually as I was running out of submissions but they generally ranged between 0.585-0.605 on the Public LB.  I mainly relied on my ensemble technique to get the score boost.\n\nThanks to the hosts and Kaggle for this interesting competition. \n\n**Inference Notebook**: https://www.kaggle.com/taggatle/cornell-birdcall-identification-1st-place-solution \n**Training Code**: https://github.com/ryanwongsa/kaggle-birdsong-recognition\n**Example on How to train the model on Kaggle Kernels**: https://www.kaggle.com/taggatle/example-training-notebook",
    "1012188": "Congrats, Ryan! Much appreciate your participation and more information on the winning solution.",
    "1012204": "Happy to know SED model won this competition! It's quite interesting that you didn't use so much data augmentation.\n\nCongratz @taggatle 🎉",
    "1012248": "Congrats on the solo win!  I like your voting ensemble, we could not find a way to ensemble our models.",
    "1012406": "simple and effective! Congrac)",
    "1012546": "Congrats @taggatle! 🎉🎉",
    "1012568": "Thanks @hidehisaarai1213 , I think you provided a lot of guidance and help with what you shared.\n\nI was a kind of afraid of adding too much augmentation as I didn't want to distort the bird calls too much. I mainly relied on the pink noise as suggested by one of the hosts.",
    "1012569": "Thanks, @holgerklinck.",
    "1012575": "Thanks, @cpmpml. Yeah I mainly used voting to reduce the false positives, I think it was the major factor in my solution since individually the models were in the silver region.",
    "1012577": "Thanks, @kupchanski",
    "1012579": "Thanks, @ramarlina",
    "1012688": "Congats @taggatle",
    "1012694": "Congratulation on the win !",
    "1012730": "I noticed that score of blend was the blend of scores, but I didn't conclude that averaging fold predictions wasn't great.  I will certainly try voting next time!",
    "1012906": "Congratulations! This will be my learning material in this weekend!",
    "1014919": "Congratulation on the win !",
    "1018815": "Congratulations to  the 1st place @taggatle!🎉",
    "1019310": "congratulations Ryan",
    "1020099": "Thank you for sharing this.",
    "1020231": "interesting and congratulations",
    "1020683": "Congratulations!!!",
    "1037838": "Could you tell me what were the scores of single models you used for ensembling? I want to know how much did hard voting contribute to your score.",
    "1038194": "I don't have all of them as I was limited on submissions but I have an example of 3 of the results on fold 0,1,2 in the attached image. At the top contains an ensemble of 5 of the folds with 3 votes. I didn't trust the LB completely so didn't want to pick individual models based on the LB score and instead based it on the ensemble score. \n\nAn ensemble of 9 models with 3 votes got private LB 0.676 and public LB 0.613, so just adding 4 more models with the same vote count scored a bit better on the private and public LB. An ensemble of 13 models with the same votes also improves the private LB to 0.679. \n\nIt is possible to get a high scoring single model though. One of my earlier models when I was hyperparameter testing got a score of 0.605 on the public LB and 0.674 on the private LB. It wasn't included in the final ensemble.",
    "1043202": "Congrats on the win!\nThanks for the clear write-up :)",
    "1043802": "Congratulation on the win !why you switch your models?   We could not find a way to use.",
    "1044666": "good one using the pre-trained model",
    "1047056": "Sorry for not replying...\nSeems it's private score (also public score) varies a lot - maybe it helped the result to be good in terms of diversity.\nThanks for telling that!",
    "1086453": "Hi there real appreciate you sharing all you work, I was wondering if you can help me.\n\ntrying to run 'train.py' but every time I get to : parser.add_argument('--config', type=str)\n\nI get this error, as the arguments are None...\n\n _StoreAction(option_strings=['--config'], dest='config', nargs=None, const=None, default=None, type=<class 'str'>, choices=None, help=None, metavar=None)\n\nThanks!",
    "1086528": "Hi, if you create a file for hyperparameters like the one in the Kaggle notebook (https://www.kaggle.com/taggatle/example-training-notebook) and save it somewhere as \"config_params/example_config.py\". Then you can use those configs for training with the following command:\n```\npython sed_train.py --config \"config_params.example_config\"\n```",
    "1086539": "Hi thanks for your replay!! really appreciate it.\n\nI was trying this but this not worked...\n\nparser.add_argument('--config', default='./config_params/configs.py', type=str)\n\nNow I see how, thanks!!",
    "1142667": "Thanks so much for sharing your great work! Wondering, why do you think you and other teams did not make use of the learnable wavegram component of the PANN paper?",
    "1247882": "taggatle very late here but any chance you could explain the intuition behind your custom loss function?",
    "1254584": "Thanks for sharing as we use this to identify birds in our backyard.",
    "1306096": "taggatle will spectogram look different to model (hence difference result) if train  on more duration clips chunk and infer on less duration clip chunk say 5 sec clip chunk"
  },
  "source": "meta"
}