{
  "id": 327047,
  "title": "1st place solution models (it’s not all BirdNet)",
  "url": "/competitions/birdclef-2022/discussion/327047",
  "author_name": "Volodymyr",
  "post_date": "2022-05-25T10:37:01.993000",
  "votes": 63,
  "comment_count": 4,
  "views": 0,
  "content": "<p>First of all, thanks to Kaggle Team, <a href=\"https://www.kaggle.com/stefankahl\" target=\"_blank\">@stefankahl</a> and all participants </p>\n<p>As you may know, in our final blend we have used BirdNet along with other models provided by me, <a href=\"https://www.kaggle.com/ivanpan\" target=\"_blank\">@ivanpan</a>, <a href=\"https://www.kaggle.com/realsleim\" target=\"_blank\">@realsleim</a> and <a href=\"https://www.kaggle.com/selimsef\" target=\"_blank\">@selimsef</a></p>\n<p>In this post I will cover my models and approach</p>\n<p><strong>Model Architecture</strong></p>\n<p>I was using SED architecture, proposed and used by <a href=\"https://www.kaggle.com/tattaka\" target=\"_blank\">@tattaka</a> - <a href=\"https://www.kaggle.com/competitions/birdclef-2021/discussion/243293\" target=\"_blank\">post</a></p>\n<p>As a backbones I have used:</p>\n<ul>\n<li>tf_efficientnet_b3_ns</li>\n<li>eca_nfnet_l0</li>\n</ul>\n<p>I have changed the first stride from (2,2) to (1,1) in order to have larger (in terms of length and number of frequencies )  output of CNN encoder ( I have taken this trick from <a href=\"https://www.kaggle.com/ilu000\" target=\"_blank\">@ilu000</a> <a href=\"https://www.kaggle.com/competitions/seti-breakthrough-listen/discussion/266385\" target=\"_blank\">SETI</a>)</p>\n<p><strong>Model Training</strong></p>\n<p>I have trained model on 15 sec chunks and with secondary labels</p>\n<p>I have used following augmentations:</p>\n<ul>\n<li>GaussianNoise</li>\n<li>PinkNoise</li>\n<li>OR Mixup on waveforms</li>\n<li>BackgroundNoise. For training - <a href=\"https://www.kaggle.com/datasets/mmoreaux/environmental-sound-classification-50\" target=\"_blank\">this dataset</a>. For finetuning - esc50 + nocall from soundscapes of 2021 BirdClef Comp</li>\n</ul>\n<p>Proposed by <a href=\"https://www.kaggle.com/selimsef\" target=\"_blank\">@selimsef</a> I have used weights (computed by <code>primary_label</code>) for Dataloader and Loss in order to cope with unbalanced dataset (especially for <code>scored_birds</code>) </p>\n<p>As for Loss I have used simple BCE on <code>clipwise</code> logits</p>\n<p>I was tracking 3 best checkpoints by LB metric and Validation loss and then averaged 3 model weights (kind of naive SWA)</p>\n<p><strong>Training stages</strong> </p>\n<p>For training I have used 2 stage training:</p>\n<ol>\n<li>Pretrain on 2021 and 2022 comp data   </li>\n<li>Finetune on data from pretrain BUT filtered by the next rule - <code>Take samples which contain scored_bird in primary_label OR secondary_labels</code></li>\n</ol>\n<p><strong>Inference</strong></p>\n<p>Having SED model I have tried 2 options for inference:</p>\n<ol>\n<li>Proposed by <a href=\"https://www.kaggle.com/tattaka\" target=\"_blank\">@tattaka</a> - using AND rule for <code>long</code> and <code>short</code> clipwise predictions. short prediction - 5 sec, long - 15 sec</li>\n<li>Feed model 15 sec chunk BUT apply head only on centered 5 sec reduced CNN image and use max(framewise, dim=time)</li>\n</ol>\n<p>Overall second option worked better for me</p>\n<p>Choosing threshold. Here I have tried also 2 options:</p>\n<ol>\n<li>Use ordinary threshold. Optimal values varied for me from 0.2 - 0.3</li>\n<li>Use quantile threshold, originally proposed by <a href=\"https://www.kaggle.com/philippsinger\" target=\"_blank\">@philippsinger</a> <a href=\"https://www.kaggle.com/competitions/birdclef-2021/discussion/243463\" target=\"_blank\">post</a>. Optimal value for <code>quantile_tresh</code> was 0.25</li>\n</ol>\n<pre><code>tresh = np.quantile(\n                        test_model_probs[:, scored_bird_ids].flatten(),\n                        1 - quantile_tresh,\n                    )\n</code></pre>\n<p>For solo model (5 folds) second worked better and for ensembles first worked better</p>\n<p><strong>Validation</strong></p>\n<p>I have used 5 CV Stratified training (also <code>maupar</code> sample was splitted on 5 samples in order to have consistent OOF score)<br>\nCompute LB metric using next prediction scheme:<br>\nFor each validation sample - slice it on pieces -&gt; predict each piece -&gt; max(sample_predictions, dim=pieces)<br>\nAnd then compute metric using <code>all_labels=[primary_label] + secondary_labels</code><br>\nAlso I have computed this metric only on samples which contain scored_bird in all_labels and taking into account only scored_birds<br>\nAnd finally optimize threshold with step 0.1 </p>\n<p><strong>Results</strong></p>\n<ul>\n<li>tf_efficientnet_b3_ns: Val = 0.87894692; Public = 0.82; Private = 0.78</li>\n<li>eca_nfnet_l0: Val = 0.88640918; Public = 0.82; Private = 0.78</li>\n</ul>\n<p><strong>Inference Kernel</strong> - <a href=\"https://www.kaggle.com/code/ivanpan/fork-of-fork-of-cls-exp-1-870246-021187-967146/notebook?scriptVersionId=96433080\" target=\"_blank\">https://www.kaggle.com/code/ivanpan/fork-of-fork-of-cls-exp-1-870246-021187-967146/notebook?scriptVersionId=96433080</a><br>\n<strong>GitHub Repo</strong> - <a href=\"https://github.com/Selimonder/birdclef-2022\" target=\"_blank\">https://github.com/Selimonder/birdclef-2022</a></p>",
  "messages": [
    {
      "id": 1800976,
      "postDate": "2022-05-25T10:37:01.993Z",
      "content": "<p>First of all, thanks to Kaggle Team, <a href=\"https://www.kaggle.com/stefankahl\" target=\"_blank\">@stefankahl</a> and all participants </p>\n<p>As you may know, in our final blend we have used BirdNet along with other models provided by me, <a href=\"https://www.kaggle.com/ivanpan\" target=\"_blank\">@ivanpan</a>, <a href=\"https://www.kaggle.com/realsleim\" target=\"_blank\">@realsleim</a> and <a href=\"https://www.kaggle.com/selimsef\" target=\"_blank\">@selimsef</a></p>\n<p>In this post I will cover my models and approach</p>\n<p><strong>Model Architecture</strong></p>\n<p>I was using SED architecture, proposed and used by <a href=\"https://www.kaggle.com/tattaka\" target=\"_blank\">@tattaka</a> - <a href=\"https://www.kaggle.com/competitions/birdclef-2021/discussion/243293\" target=\"_blank\">post</a></p>\n<p>As a backbones I have used:</p>\n<ul>\n<li>tf_efficientnet_b3_ns</li>\n<li>eca_nfnet_l0</li>\n</ul>\n<p>I have changed the first stride from (2,2) to (1,1) in order to have larger (in terms of length and number of frequencies )  output of CNN encoder ( I have taken this trick from <a href=\"https://www.kaggle.com/ilu000\" target=\"_blank\">@ilu000</a> <a href=\"https://www.kaggle.com/competitions/seti-breakthrough-listen/discussion/266385\" target=\"_blank\">SETI</a>)</p>\n<p><strong>Model Training</strong></p>\n<p>I have trained model on 15 sec chunks and with secondary labels</p>\n<p>I have used following augmentations:</p>\n<ul>\n<li>GaussianNoise</li>\n<li>PinkNoise</li>\n<li>OR Mixup on waveforms</li>\n<li>BackgroundNoise. For training - <a href=\"https://www.kaggle.com/datasets/mmoreaux/environmental-sound-classification-50\" target=\"_blank\">this dataset</a>. For finetuning - esc50 + nocall from soundscapes of 2021 BirdClef Comp</li>\n</ul>\n<p>Proposed by <a href=\"https://www.kaggle.com/selimsef\" target=\"_blank\">@selimsef</a> I have used weights (computed by <code>primary_label</code>) for Dataloader and Loss in order to cope with unbalanced dataset (especially for <code>scored_birds</code>) </p>\n<p>As for Loss I have used simple BCE on <code>clipwise</code> logits</p>\n<p>I was tracking 3 best checkpoints by LB metric and Validation loss and then averaged 3 model weights (kind of naive SWA)</p>\n<p><strong>Training stages</strong> </p>\n<p>For training I have used 2 stage training:</p>\n<ol>\n<li>Pretrain on 2021 and 2022 comp data   </li>\n<li>Finetune on data from pretrain BUT filtered by the next rule - <code>Take samples which contain scored_bird in primary_label OR secondary_labels</code></li>\n</ol>\n<p><strong>Inference</strong></p>\n<p>Having SED model I have tried 2 options for inference:</p>\n<ol>\n<li>Proposed by <a href=\"https://www.kaggle.com/tattaka\" target=\"_blank\">@tattaka</a> - using AND rule for <code>long</code> and <code>short</code> clipwise predictions. short prediction - 5 sec, long - 15 sec</li>\n<li>Feed model 15 sec chunk BUT apply head only on centered 5 sec reduced CNN image and use max(framewise, dim=time)</li>\n</ol>\n<p>Overall second option worked better for me</p>\n<p>Choosing threshold. Here I have tried also 2 options:</p>\n<ol>\n<li>Use ordinary threshold. Optimal values varied for me from 0.2 - 0.3</li>\n<li>Use quantile threshold, originally proposed by <a href=\"https://www.kaggle.com/philippsinger\" target=\"_blank\">@philippsinger</a> <a href=\"https://www.kaggle.com/competitions/birdclef-2021/discussion/243463\" target=\"_blank\">post</a>. Optimal value for <code>quantile_tresh</code> was 0.25</li>\n</ol>\n<pre><code>tresh = np.quantile(\n                        test_model_probs[:, scored_bird_ids].flatten(),\n                        1 - quantile_tresh,\n                    )\n</code></pre>\n<p>For solo model (5 folds) second worked better and for ensembles first worked better</p>\n<p><strong>Validation</strong></p>\n<p>I have used 5 CV Stratified training (also <code>maupar</code> sample was splitted on 5 samples in order to have consistent OOF score)<br>\nCompute LB metric using next prediction scheme:<br>\nFor each validation sample - slice it on pieces -&gt; predict each piece -&gt; max(sample_predictions, dim=pieces)<br>\nAnd then compute metric using <code>all_labels=[primary_label] + secondary_labels</code><br>\nAlso I have computed this metric only on samples which contain scored_bird in all_labels and taking into account only scored_birds<br>\nAnd finally optimize threshold with step 0.1 </p>\n<p><strong>Results</strong></p>\n<ul>\n<li>tf_efficientnet_b3_ns: Val = 0.87894692; Public = 0.82; Private = 0.78</li>\n<li>eca_nfnet_l0: Val = 0.88640918; Public = 0.82; Private = 0.78</li>\n</ul>\n<p><strong>Inference Kernel</strong> - <a href=\"https://www.kaggle.com/code/ivanpan/fork-of-fork-of-cls-exp-1-870246-021187-967146/notebook?scriptVersionId=96433080\" target=\"_blank\">https://www.kaggle.com/code/ivanpan/fork-of-fork-of-cls-exp-1-870246-021187-967146/notebook?scriptVersionId=96433080</a><br>\n<strong>GitHub Repo</strong> - <a href=\"https://github.com/Selimonder/birdclef-2022\" target=\"_blank\">https://github.com/Selimonder/birdclef-2022</a></p>",
      "rawMarkdown": "First of all, thanks to Kaggle Team, @stefankahl and all participants \n\nAs you may know, in our final blend we have used BirdNet along with other models provided by me, @ivanpan, @realsleim and @selimsef\n\nIn this post I will cover my models and approach\n\n**Model Architecture**\n\nI was using SED architecture, proposed and used by @tattaka - [post](https://www.kaggle.com/competitions/birdclef-2021/discussion/243293)\n\nAs a backbones I have used:\n- tf_efficientnet_b3_ns\n- eca_nfnet_l0\n\nI have changed the first stride from (2,2) to (1,1) in order to have larger (in terms of length and number of frequencies )  output of CNN encoder ( I have taken this trick from @ilu000 [SETI](https://www.kaggle.com/competitions/seti-breakthrough-listen/discussion/266385))\n\n**Model Training**\n\nI have trained model on 15 sec chunks and with secondary labels\n\nI have used following augmentations:\n- GaussianNoise\n- PinkNoise\n- OR Mixup on waveforms\n- BackgroundNoise. For training - [this dataset](https://www.kaggle.com/datasets/mmoreaux/environmental-sound-classification-50). For finetuning - esc50 + nocall from soundscapes of 2021 BirdClef Comp\n\nProposed by @selimsef I have used weights (computed by `primary_label`) for Dataloader and Loss in order to cope with unbalanced dataset (especially for `scored_birds`) \n\nAs for Loss I have used simple BCE on `clipwise` logits\n\nI was tracking 3 best checkpoints by LB metric and Validation loss and then averaged 3 model weights (kind of naive SWA)\n\n**Training stages** \n\nFor training I have used 2 stage training:\n1. Pretrain on 2021 and 2022 comp data   \n2. Finetune on data from pretrain BUT filtered by the next rule - `Take samples which contain scored_bird in primary_label OR secondary_labels`\n\n**Inference**\n\nHaving SED model I have tried 2 options for inference:\n1. Proposed by @tattaka - using AND rule for `long` and `short` clipwise predictions. short prediction - 5 sec, long - 15 sec\n2. Feed model 15 sec chunk BUT apply head only on centered 5 sec reduced CNN image and use max(framewise, dim=time)\n\nOverall second option worked better for me\n\nChoosing threshold. Here I have tried also 2 options:\n1. Use ordinary threshold. Optimal values varied for me from 0.2 - 0.3\n2. Use quantile threshold, originally proposed by @philippsinger [post](https://www.kaggle.com/competitions/birdclef-2021/discussion/243463). Optimal value for `quantile_tresh` was 0.25\n```\ntresh = np.quantile(\n                        test_model_probs[:, scored_bird_ids].flatten(),\n                        1 - quantile_tresh,\n                    )\n```\n\nFor solo model (5 folds) second worked better and for ensembles first worked better\n\n**Validation**\n\nI have used 5 CV Stratified training (also `maupar` sample was splitted on 5 samples in order to have consistent OOF score)\nCompute LB metric using next prediction scheme:\nFor each validation sample - slice it on pieces -> predict each piece -> max(sample_predictions, dim=pieces)\nAnd then compute metric using `all_labels=[primary_label] + secondary_labels`\nAlso I have computed this metric only on samples which contain scored_bird in all_labels and taking into account only scored_birds\nAnd finally optimize threshold with step 0.1 \n\n**Results**\n\n- tf_efficientnet_b3_ns: Val = 0.87894692; Public = 0.82; Private = 0.78\n- eca_nfnet_l0: Val = 0.88640918; Public = 0.82; Private = 0.78\n\n**Inference Kernel** - https://www.kaggle.com/code/ivanpan/fork-of-fork-of-cls-exp-1-870246-021187-967146/notebook?scriptVersionId=96433080\n**GitHub Repo** - https://github.com/Selimonder/birdclef-2022",
      "votes": 63
    },
    {
      "id": 1801212,
      "postDate": "2022-05-25T14:27:50.513Z",
      "content": "<p>Congrats on 1st place! Thanks for sharing your solution. Amazing work!</p>",
      "rawMarkdown": "Congrats on 1st place! Thanks for sharing your solution. Amazing work!",
      "votes": 1
    },
    {
      "id": 1814478,
      "postDate": "2022-06-08T00:44:46.790Z",
      "content": "<p>Thanks a lot, the solution description and code repo!</p>",
      "rawMarkdown": "Thanks a lot, the solution description and code repo!"
    },
    {
      "id": 1806622,
      "postDate": "2022-05-31T10:43:22.253Z",
      "content": "<p>Congratulations for the 1st place, <a href=\"https://www.kaggle.com/vladimirsydor\" target=\"_blank\">@vladimirsydor</a>. Regarding your 2 stages training, in your 2nd finetune stage training, di d you train all layers of your model, or freeze backbone to avoid overfitting? </p>\n<blockquote>\n  <p>Finetune on data from pretrain BUT filtered by the next rule - Take samples which contain scored_bird in primary_label OR secondary_labels</p>\n</blockquote>",
      "rawMarkdown": "Congratulations for the 1st place, @vladimirsydor. Regarding your 2 stages training, in your 2nd finetune stage training, di d you train all layers of your model, or freeze backbone to avoid overfitting? \n\n> Finetune on data from pretrain BUT filtered by the next rule - Take samples which contain scored_bird in primary_label OR secondary_labels"
    },
    {
      "id": 2211880,
      "postDate": "2023-04-06T10:36:44.313Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 1801212,
      "author_name": "Takashi Tamura",
      "author_url": "",
      "post_date": "2022-05-25T14:27:50.513000",
      "content": "<p>Congrats on 1st place! Thanks for sharing your solution. Amazing work!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1814478,
      "author_name": "Lainey",
      "author_url": "",
      "post_date": "2022-06-08T00:44:46.790000",
      "content": "<p>Thanks a lot, the solution description and code repo!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1806622,
      "author_name": "Hao",
      "author_url": "",
      "post_date": "2022-05-31T10:43:22.253000",
      "content": "<p>Congratulations for the 1st place, <a href=\"https://www.kaggle.com/vladimirsydor\" target=\"_blank\">@vladimirsydor</a>. Regarding your 2 stages training, in your 2nd finetune stage training, di d you train all layers of your model, or freeze backbone to avoid overfitting? </p>\n<blockquote>\n  <p>Finetune on data from pretrain BUT filtered by the next rule - Take samples which contain scored_bird in primary_label OR secondary_labels</p>\n</blockquote>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2211880,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-04-06T10:36:44.313000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1800976": "First of all, thanks to Kaggle Team, @stefankahl and all participants \n\nAs you may know, in our final blend we have used BirdNet along with other models provided by me, @ivanpan, @realsleim and @selimsef\n\nIn this post I will cover my models and approach\n\n**Model Architecture**\n\nI was using SED architecture, proposed and used by @tattaka - [post](https://www.kaggle.com/competitions/birdclef-2021/discussion/243293)\n\nAs a backbones I have used:\n- tf_efficientnet_b3_ns\n- eca_nfnet_l0\n\nI have changed the first stride from (2,2) to (1,1) in order to have larger (in terms of length and number of frequencies )  output of CNN encoder ( I have taken this trick from @ilu000 [SETI](https://www.kaggle.com/competitions/seti-breakthrough-listen/discussion/266385))\n\n**Model Training**\n\nI have trained model on 15 sec chunks and with secondary labels\n\nI have used following augmentations:\n- GaussianNoise\n- PinkNoise\n- OR Mixup on waveforms\n- BackgroundNoise. For training - [this dataset](https://www.kaggle.com/datasets/mmoreaux/environmental-sound-classification-50). For finetuning - esc50 + nocall from soundscapes of 2021 BirdClef Comp\n\nProposed by @selimsef I have used weights (computed by `primary_label`) for Dataloader and Loss in order to cope with unbalanced dataset (especially for `scored_birds`) \n\nAs for Loss I have used simple BCE on `clipwise` logits\n\nI was tracking 3 best checkpoints by LB metric and Validation loss and then averaged 3 model weights (kind of naive SWA)\n\n**Training stages** \n\nFor training I have used 2 stage training:\n1. Pretrain on 2021 and 2022 comp data   \n2. Finetune on data from pretrain BUT filtered by the next rule - `Take samples which contain scored_bird in primary_label OR secondary_labels`\n\n**Inference**\n\nHaving SED model I have tried 2 options for inference:\n1. Proposed by @tattaka - using AND rule for `long` and `short` clipwise predictions. short prediction - 5 sec, long - 15 sec\n2. Feed model 15 sec chunk BUT apply head only on centered 5 sec reduced CNN image and use max(framewise, dim=time)\n\nOverall second option worked better for me\n\nChoosing threshold. Here I have tried also 2 options:\n1. Use ordinary threshold. Optimal values varied for me from 0.2 - 0.3\n2. Use quantile threshold, originally proposed by @philippsinger [post](https://www.kaggle.com/competitions/birdclef-2021/discussion/243463). Optimal value for `quantile_tresh` was 0.25\n```\ntresh = np.quantile(\n                        test_model_probs[:, scored_bird_ids].flatten(),\n                        1 - quantile_tresh,\n                    )\n```\n\nFor solo model (5 folds) second worked better and for ensembles first worked better\n\n**Validation**\n\nI have used 5 CV Stratified training (also `maupar` sample was splitted on 5 samples in order to have consistent OOF score)\nCompute LB metric using next prediction scheme:\nFor each validation sample - slice it on pieces -> predict each piece -> max(sample_predictions, dim=pieces)\nAnd then compute metric using `all_labels=[primary_label] + secondary_labels`\nAlso I have computed this metric only on samples which contain scored_bird in all_labels and taking into account only scored_birds\nAnd finally optimize threshold with step 0.1 \n\n**Results**\n\n- tf_efficientnet_b3_ns: Val = 0.87894692; Public = 0.82; Private = 0.78\n- eca_nfnet_l0: Val = 0.88640918; Public = 0.82; Private = 0.78\n\n**Inference Kernel** - https://www.kaggle.com/code/ivanpan/fork-of-fork-of-cls-exp-1-870246-021187-967146/notebook?scriptVersionId=96433080\n**GitHub Repo** - https://github.com/Selimonder/birdclef-2022",
    "1801212": "Congrats on 1st place! Thanks for sharing your solution. Amazing work!",
    "1814478": "Thanks a lot, the solution description and code repo!",
    "1806622": "Congratulations for the 1st place, @vladimirsydor. Regarding your 2 stages training, in your 2nd finetune stage training, di d you train all layers of your model, or freeze backbone to avoid overfitting? \n\n> Finetune on data from pretrain BUT filtered by the next rule - Take samples which contain scored_bird in primary_label OR secondary_labels",
    "2211880": ""
  }
}