{
  "id": 220972,
  "title": "23rd Place Solution: Supervised Contrastive Learning Meets Domain Generalization (with TF code)",
  "url": "/competitions/rfcx-species-audio-detection/writeups/some-team-name-23rd-place-solution-supervised-cont",
  "author_name": "",
  "post_date": "2021-02-20T10:45:01.723Z",
  "votes": 29,
  "comment_count": 4,
  "views": 0,
  "content": "<h3>Introduction</h3>\n<p>Thanks Kaggle for this exciting competition and our team ( <a href=\"https://www.kaggle.com/dathudeptrai\" target=\"_blank\">@dathudeptrai</a> <a href=\"https://www.kaggle.com/mcggood\" target=\"_blank\">@mcggood</a> <a href=\"https://www.kaggle.com/akensert\" target=\"_blank\">@akensert</a> <a href=\"https://www.kaggle.com/theoviel\" target=\"_blank\">@theoviel</a> <a href=\"https://www.kaggle.com/ratthachat\" target=\"_blank\">@ratthachat</a> ) congratulate all winners -- we have learned a lot from this competition and all winners’ solutions !!</p>\n<p>From the winner solutions, it turns out there are mainly 4 ways to break 0.95x</p>\n<ol>\n<li>Masked loss</li>\n<li>Post-processing</li>\n<li>Pseudo-labelling</li>\n<li>Extra labeling</li>\n</ol>\n<p>We actually tried the first three but unfortunately could not really make them work effectively.<br>\nHere, as an alternative solution, we would love to share our own solution which is able to reach 0.943 Private. </p>\n<h3>Training pipeline (improve from 0.80x baseline to 0.92x)</h3>\n<h4>Baseline</h4>\n<p>Our baseline models score slightly above 0.8. We adopt an audio tagging approach using a densenet121 backbone and the BCE loss. </p>\n<p>Our training pipeline includes the following tricks :</p>\n<ul>\n<li>Class-balanced sampling, using 8 classes per batch. Batch sizes were usually 64, and 32 for bigger models.</li>\n<li>Cyclical learning rate with min_lr is 0.0001 and max_lr is 0.001, step_size is 3 epochs. We train models for 100 epochs with early stopping. </li>\n<li>LookAhead with Adam optimizer (sync_period is 10 and slow_step_size is 0.5)</li>\n</ul>\n<h4>Pretraining with Supervised contrastive Learning (SCL) [0.81x -&gt; 0.85x]</h4>\n<p>Because of the small amount of data, models overfit quickly. To solve this problem,two options were using external data and cleverly pretraining our models. Unlike a lot of competitors, we focused on self-pretraining techniques : <a href=\"https://www.kaggle.com/dathudeptrai\" target=\"_blank\">@dathudeptrai</a> tried auto-encoders, GANs, SimCLR, Cola, and <a href=\"https://arxiv.org/abs/2004.11362\" target=\"_blank\">Supervised Contrastive Learning</a> which ultimately was the only thing to work. </p>\n<h4>Non-overlap time Cutmix [0.85x -&gt; 0.88x]</h4>\n<p><img src=\"https://i.imgur.com/gH2ONQh.png\" alt=\"\"></p>\n<p>Our sampling strategy consists of randomly selecting a crop containing the label. Most of the time, crops are bigger than labels which introduces false positives. One idea to make full use of our windows was to adapt cutmix to concatenate samples such that labels are entirely kept (when possible). </p>\n<h4>Domain Generalization with MixStyle [0.88x -&gt; 0.89x]</h4>\n<p>Domain shift always exists in deep learning, in both practice and kaggle challenges, especially for small data. Therefore, domain generalization techniques should help with robustness. We applied a simple yet effective technique called <a href=\"https://openreview.net/pdf?id=6xHJ37MVxxp\" target=\"_blank\">Mixstyle</a>.</p>\n<h4>Multi Scale inference (MSI) [0.89x -&gt; 0.91x]</h4>\n<p>Duration of species’ call varies quite a lot. For example, for class 3 it is around 0.7 seconds while for class 23 is around 8 seconds. To use this prior information, we use multiple window sizes (instead of using a single one). For each class, we choose the one that yields the best CV. In case we have multiple window sizes reaching the maximum, we take the largest window. Although our CV setup which consists of crops centered around the labels did not correlate really well with LB, the 2% CV improvement reflected on LB quite well.</p>\n<h4>Positive learning and Negative learning  [0.91x -&gt; 0.92x]</h4>\n<p>We used the following assumption to improve the training of our models : <br>\nFor a given recording, if a species has a FP and no TP, then it is not in the call. Our BCE was then updated to make sure the model predicts 0 for such species. </p>\n<h3>Ensembling</h3>\n<p>Our best single model densenet121 scores around 0.92 public and 0.93 private. Averaging some models with different backbones, we were able to reach 0.937. We tried many different ensembling, scale fixing and post-processing ideas, and were able to improve our score a bit, but unfortunately we could not find the real magic.</p>\n<p>In the end, we empirically analyzed the most uncertain class predictions from our best models, and averaged predictions with other (weaker) models. We relied on diversity to make our submission more robust. Our final ensemble scored public 0.942 and private 0.943.</p>\n<h5>Thanks for reading !</h5>\n<p><a href=\"https://github.com/dathudeptrai/rfcx-kaggle\" target=\"_blank\">TensorFlow Code Here</a></p>",
  "messages": [
    {
      "id": "1211554",
      "postDate": "02/20/2021 10:17:02",
      "content": "<h3>Introduction</h3>\n<p>Thanks Kaggle for this exciting competition and our team ( <a href=\"https://www.kaggle.com/dathudeptrai\" target=\"_blank\">@dathudeptrai</a> <a href=\"https://www.kaggle.com/mcggood\" target=\"_blank\">@mcggood</a> <a href=\"https://www.kaggle.com/akensert\" target=\"_blank\">@akensert</a> <a href=\"https://www.kaggle.com/theoviel\" target=\"_blank\">@theoviel</a> <a href=\"https://www.kaggle.com/ratthachat\" target=\"_blank\">@ratthachat</a> ) congratulate all winners -- we have learned a lot from this competition and all winners’ solutions !!</p>\n<p>From the winner solutions, it turns out there are mainly 4 ways to break 0.95x</p>\n<ol>\n<li>Masked loss</li>\n<li>Post-processing</li>\n<li>Pseudo-labelling</li>\n<li>Extra labeling</li>\n</ol>\n<p>We actually tried the first three but unfortunately could not really make them work effectively.<br>\nHere, as an alternative solution, we would love to share our own solution which is able to reach 0.943 Private. </p>\n<h3>Training pipeline (improve from 0.80x baseline to 0.92x)</h3>\n<h4>Baseline</h4>\n<p>Our baseline models score slightly above 0.8. We adopt an audio tagging approach using a densenet121 backbone and the BCE loss. </p>\n<p>Our training pipeline includes the following tricks :</p>\n<ul>\n<li>Class-balanced sampling, using 8 classes per batch. Batch sizes were usually 64, and 32 for bigger models.</li>\n<li>Cyclical learning rate with min_lr is 0.0001 and max_lr is 0.001, step_size is 3 epochs. We train models for 100 epochs with early stopping. </li>\n<li>LookAhead with Adam optimizer (sync_period is 10 and slow_step_size is 0.5)</li>\n</ul>\n<h4>Pretraining with Supervised contrastive Learning (SCL) [0.81x -&gt; 0.85x]</h4>\n<p>Because of the small amount of data, models overfit quickly. To solve this problem,two options were using external data and cleverly pretraining our models. Unlike a lot of competitors, we focused on self-pretraining techniques : <a href=\"https://www.kaggle.com/dathudeptrai\" target=\"_blank\">@dathudeptrai</a> tried auto-encoders, GANs, SimCLR, Cola, and <a href=\"https://arxiv.org/abs/2004.11362\" target=\"_blank\">Supervised Contrastive Learning</a> which ultimately was the only thing to work. </p>\n<h4>Non-overlap time Cutmix [0.85x -&gt; 0.88x]</h4>\n<p><img src=\"https://i.imgur.com/gH2ONQh.png\" alt=\"\"></p>\n<p>Our sampling strategy consists of randomly selecting a crop containing the label. Most of the time, crops are bigger than labels which introduces false positives. One idea to make full use of our windows was to adapt cutmix to concatenate samples such that labels are entirely kept (when possible). </p>\n<h4>Domain Generalization with MixStyle [0.88x -&gt; 0.89x]</h4>\n<p>Domain shift always exists in deep learning, in both practice and kaggle challenges, especially for small data. Therefore, domain generalization techniques should help with robustness. We applied a simple yet effective technique called <a href=\"https://openreview.net/pdf?id=6xHJ37MVxxp\" target=\"_blank\">Mixstyle</a>.</p>\n<h4>Multi Scale inference (MSI) [0.89x -&gt; 0.91x]</h4>\n<p>Duration of species’ call varies quite a lot. For example, for class 3 it is around 0.7 seconds while for class 23 is around 8 seconds. To use this prior information, we use multiple window sizes (instead of using a single one). For each class, we choose the one that yields the best CV. In case we have multiple window sizes reaching the maximum, we take the largest window. Although our CV setup which consists of crops centered around the labels did not correlate really well with LB, the 2% CV improvement reflected on LB quite well.</p>\n<h4>Positive learning and Negative learning  [0.91x -&gt; 0.92x]</h4>\n<p>We used the following assumption to improve the training of our models : <br>\nFor a given recording, if a species has a FP and no TP, then it is not in the call. Our BCE was then updated to make sure the model predicts 0 for such species. </p>\n<h3>Ensembling</h3>\n<p>Our best single model densenet121 scores around 0.92 public and 0.93 private. Averaging some models with different backbones, we were able to reach 0.937. We tried many different ensembling, scale fixing and post-processing ideas, and were able to improve our score a bit, but unfortunately we could not find the real magic.</p>\n<p>In the end, we empirically analyzed the most uncertain class predictions from our best models, and averaged predictions with other (weaker) models. We relied on diversity to make our submission more robust. Our final ensemble scored public 0.942 and private 0.943.</p>\n<h5>Thanks for reading !</h5>\n<p><a href=\"https://github.com/dathudeptrai/rfcx-kaggle\" target=\"_blank\">TensorFlow Code Here</a></p>",
      "rawMarkdown": "### Introduction\n\nThanks Kaggle for this exciting competition and our team ( @dathudeptrai @mcggood @akensert @theoviel @ratthachat ) congratulate all winners -- we have learned a lot from this competition and all winners’ solutions !!\n\nFrom the winner solutions, it turns out there are mainly 4 ways to break 0.95x\n\n1. Masked loss\n2. Post-processing\n3. Pseudo-labelling\n4. Extra labeling\n\n\nWe actually tried the first three but unfortunately could not really make them work effectively.\nHere, as an alternative solution, we would love to share our own solution which is able to reach 0.943 Private. \n\n### Training pipeline (improve from 0.80x baseline to 0.92x)\n\n#### Baseline\n\nOur baseline models score slightly above 0.8. We adopt an audio tagging approach using a densenet121 backbone and the BCE loss. \n\nOur training pipeline includes the following tricks :\n \n- Class-balanced sampling, using 8 classes per batch. Batch sizes were usually 64, and 32 for bigger models.\n- Cyclical learning rate with min_lr is 0.0001 and max_lr is 0.001, step_size is 3 epochs. We train models for 100 epochs with early stopping. \n- LookAhead with Adam optimizer (sync_period is 10 and slow_step_size is 0.5)\n\n#### Pretraining with Supervised contrastive Learning (SCL) [0.81x -> 0.85x]\n\nBecause of the small amount of data, models overfit quickly. To solve this problem,two options were using external data and cleverly pretraining our models. Unlike a lot of competitors, we focused on self-pretraining techniques : @dathudeptrai tried auto-encoders, GANs, SimCLR, Cola, and [Supervised Contrastive Learning](https://arxiv.org/abs/2004.11362) which ultimately was the only thing to work. \n\n\n#### Non-overlap time Cutmix [0.85x -> 0.88x]\n\n![](https://i.imgur.com/gH2ONQh.png)\n\nOur sampling strategy consists of randomly selecting a crop containing the label. Most of the time, crops are bigger than labels which introduces false positives. One idea to make full use of our windows was to adapt cutmix to concatenate samples such that labels are entirely kept (when possible). \n\n#### Domain Generalization with MixStyle [0.88x -> 0.89x]\n\nDomain shift always exists in deep learning, in both practice and kaggle challenges, especially for small data. Therefore, domain generalization techniques should help with robustness. We applied a simple yet effective technique called [Mixstyle](https://openreview.net/pdf?id=6xHJ37MVxxp).\n\n\n#### Multi Scale inference (MSI) [0.89x -> 0.91x]\n\nDuration of species’ call varies quite a lot. For example, for class 3 it is around 0.7 seconds while for class 23 is around 8 seconds. To use this prior information, we use multiple window sizes (instead of using a single one). For each class, we choose the one that yields the best CV. In case we have multiple window sizes reaching the maximum, we take the largest window. Although our CV setup which consists of crops centered around the labels did not correlate really well with LB, the 2% CV improvement reflected on LB quite well.\n\n\n#### Positive learning and Negative learning  [0.91x -> 0.92x]\n\nWe used the following assumption to improve the training of our models : \nFor a given recording, if a species has a FP and no TP, then it is not in the call. Our BCE was then updated to make sure the model predicts 0 for such species. \n\n### Ensembling\n\nOur best single model densenet121 scores around 0.92 public and 0.93 private. Averaging some models with different backbones, we were able to reach 0.937. We tried many different ensembling, scale fixing and post-processing ideas, and were able to improve our score a bit, but unfortunately we could not find the real magic.\n\nIn the end, we empirically analyzed the most uncertain class predictions from our best models, and averaged predictions with other (weaker) models. We relied on diversity to make our submission more robust. Our final ensemble scored public 0.942 and private 0.943.\n\n##### Thanks for reading ! \n\n[TensorFlow Code Here](https://github.com/dathudeptrai/rfcx-kaggle)",
      "votes": null
    },
    {
      "id": "1211556",
      "postDate": "02/20/2021 10:26:29",
      "content": "<p>Wow interesting solution! Good work and thanks for sharing <a href=\"https://www.kaggle.com/dathudeptrai\" target=\"_blank\">@dathudeptrai</a> </p>",
      "rawMarkdown": "Wow interesting solution! Good work and thanks for sharing @dathudeptrai",
      "votes": null
    },
    {
      "id": "1211561",
      "postDate": "02/20/2021 10:29:45",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/duykhanh99\" target=\"_blank\">@duykhanh99</a> !</p>",
      "rawMarkdown": "Thanks @duykhanh99 !",
      "votes": null
    },
    {
      "id": "1211576",
      "postDate": "02/20/2021 10:37:05",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/dathudeptrai\" target=\"_blank\">@dathudeptrai</a> <a href=\"https://www.kaggle.com/mcggood\" target=\"_blank\">@mcggood</a> <a href=\"https://www.kaggle.com/akensert\" target=\"_blank\">@akensert</a> <a href=\"https://www.kaggle.com/ratthachat\" target=\"_blank\">@ratthachat</a> for the nice competition, it was a pleasure !</p>",
      "rawMarkdown": "Thanks @dathudeptrai @mcggood @akensert @ratthachat for the nice competition, it was a pleasure !",
      "votes": null
    },
    {
      "id": "1219715",
      "postDate": "02/27/2021 07:00:25",
      "content": "<p>Thanks for sharing the code, I have starred!</p>",
      "rawMarkdown": "Thanks for sharing the code, I have starred!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1211556,
      "author_name": "duykhanh99",
      "author_url": "",
      "post_date": "02/20/2021 10:26:29",
      "content": "<p>Wow interesting solution! Good work and thanks for sharing <a href=\"https://www.kaggle.com/dathudeptrai\" target=\"_blank\">@dathudeptrai</a> </p>",
      "votes": null,
      "replies": [
        {
          "id": 1211561,
          "author_name": "theoviel",
          "author_url": "",
          "post_date": "02/20/2021 10:29:45",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/duykhanh99\" target=\"_blank\">@duykhanh99</a> !</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1211576,
      "author_name": "theoviel",
      "author_url": "",
      "post_date": "02/20/2021 10:37:05",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/dathudeptrai\" target=\"_blank\">@dathudeptrai</a> <a href=\"https://www.kaggle.com/mcggood\" target=\"_blank\">@mcggood</a> <a href=\"https://www.kaggle.com/akensert\" target=\"_blank\">@akensert</a> <a href=\"https://www.kaggle.com/ratthachat\" target=\"_blank\">@ratthachat</a> for the nice competition, it was a pleasure !</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1219715,
      "author_name": "wubinbai",
      "author_url": "",
      "post_date": "02/27/2021 07:00:25",
      "content": "<p>Thanks for sharing the code, I have starred!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1211554": "### Introduction\n\nThanks Kaggle for this exciting competition and our team ( @dathudeptrai @mcggood @akensert @theoviel @ratthachat ) congratulate all winners -- we have learned a lot from this competition and all winners’ solutions !!\n\nFrom the winner solutions, it turns out there are mainly 4 ways to break 0.95x\n\n1. Masked loss\n2. Post-processing\n3. Pseudo-labelling\n4. Extra labeling\n\n\nWe actually tried the first three but unfortunately could not really make them work effectively.\nHere, as an alternative solution, we would love to share our own solution which is able to reach 0.943 Private. \n\n### Training pipeline (improve from 0.80x baseline to 0.92x)\n\n#### Baseline\n\nOur baseline models score slightly above 0.8. We adopt an audio tagging approach using a densenet121 backbone and the BCE loss. \n\nOur training pipeline includes the following tricks :\n \n- Class-balanced sampling, using 8 classes per batch. Batch sizes were usually 64, and 32 for bigger models.\n- Cyclical learning rate with min_lr is 0.0001 and max_lr is 0.001, step_size is 3 epochs. We train models for 100 epochs with early stopping. \n- LookAhead with Adam optimizer (sync_period is 10 and slow_step_size is 0.5)\n\n#### Pretraining with Supervised contrastive Learning (SCL) [0.81x -> 0.85x]\n\nBecause of the small amount of data, models overfit quickly. To solve this problem,two options were using external data and cleverly pretraining our models. Unlike a lot of competitors, we focused on self-pretraining techniques : @dathudeptrai tried auto-encoders, GANs, SimCLR, Cola, and [Supervised Contrastive Learning](https://arxiv.org/abs/2004.11362) which ultimately was the only thing to work. \n\n\n#### Non-overlap time Cutmix [0.85x -> 0.88x]\n\n![](https://i.imgur.com/gH2ONQh.png)\n\nOur sampling strategy consists of randomly selecting a crop containing the label. Most of the time, crops are bigger than labels which introduces false positives. One idea to make full use of our windows was to adapt cutmix to concatenate samples such that labels are entirely kept (when possible). \n\n#### Domain Generalization with MixStyle [0.88x -> 0.89x]\n\nDomain shift always exists in deep learning, in both practice and kaggle challenges, especially for small data. Therefore, domain generalization techniques should help with robustness. We applied a simple yet effective technique called [Mixstyle](https://openreview.net/pdf?id=6xHJ37MVxxp).\n\n\n#### Multi Scale inference (MSI) [0.89x -> 0.91x]\n\nDuration of species’ call varies quite a lot. For example, for class 3 it is around 0.7 seconds while for class 23 is around 8 seconds. To use this prior information, we use multiple window sizes (instead of using a single one). For each class, we choose the one that yields the best CV. In case we have multiple window sizes reaching the maximum, we take the largest window. Although our CV setup which consists of crops centered around the labels did not correlate really well with LB, the 2% CV improvement reflected on LB quite well.\n\n\n#### Positive learning and Negative learning  [0.91x -> 0.92x]\n\nWe used the following assumption to improve the training of our models : \nFor a given recording, if a species has a FP and no TP, then it is not in the call. Our BCE was then updated to make sure the model predicts 0 for such species. \n\n### Ensembling\n\nOur best single model densenet121 scores around 0.92 public and 0.93 private. Averaging some models with different backbones, we were able to reach 0.937. We tried many different ensembling, scale fixing and post-processing ideas, and were able to improve our score a bit, but unfortunately we could not find the real magic.\n\nIn the end, we empirically analyzed the most uncertain class predictions from our best models, and averaged predictions with other (weaker) models. We relied on diversity to make our submission more robust. Our final ensemble scored public 0.942 and private 0.943.\n\n##### Thanks for reading ! \n\n[TensorFlow Code Here](https://github.com/dathudeptrai/rfcx-kaggle)",
    "1211556": "Wow interesting solution! Good work and thanks for sharing @dathudeptrai",
    "1211561": "Thanks @duykhanh99 !",
    "1211576": "Thanks @dathudeptrai @mcggood @akensert @ratthachat for the nice competition, it was a pleasure !",
    "1219715": "Thanks for sharing the code, I have starred!"
  },
  "source": "meta"
}