{
  "id": 511510,
  "title": "9th place recap",
  "url": "/competitions/birdclef-2024/writeups/aphysict-9th-place-recap",
  "author_name": "",
  "post_date": "2024-06-11T12:55:01.120Z",
  "votes": 27,
  "comment_count": 7,
  "views": 0,
  "content": "<p>Share some personal experiences in this competition.</p>\n<h2><strong>CV strategy</strong></h2>\n<p>No CV, only public leaderboard. It just works.<br>\nAt early stage, I tried to build CV and even wasted lots of submissions to do leaderboard probing. Then I realized that CV is just waste of time. I just assume unlabeled soundscapes, public test set and private test set share similar data distribution later on.</p>\n<h2><strong>Data Augmentation</strong></h2>\n<p>Only cutmix and mixup. Only one model with time &amp; freq mask and it overfits.<br>\nTraditional data augmentation tends to overfit the train data, because it teaches the model to output the same representation for data from single audio clip (single domain knowledge). While things like mixup have some cross domain knowledge, because data comes from different audio clips. (not really sure, just guessing)<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F13324082%2F5bda7d799dd9fc33b17de71030a35fbf%2FScreen%20Shot%202024-06-11%20at%2010.09.53%20AM.png?generation=1718071832807423&amp;alt=media\" alt=\"nice pic from OPTIMAL REPRESENTATIONS FOR COVARIATE SHIFT\"></p>\n<h2><strong>Some Model Design</strong></h2>\n<p>Almost last year's model. Add a bottleneck layer to let the representation more compact.<br>\n<code>self.bottleneck_layer = nn.Sequential(\n            nn.Linear(backbone_out, cfg.bottleneck_dim),\n            nn.BatchNorm1d(cfg.bottleneck_dim),\n            nn.ReLU())</code><br>\nbottleneck_dim=1024(actually 512 is better). This can improve score a little(less than 0.01).</p>\n<p>With eca_nfnet_l0 as backbone and above tips, my baseline public score is 0.65(close to 0.66).</p>\n<h2><strong>Contrastive Adversarial Domain (CAD) bottleneck</strong></h2>\n<p>from the paper <a href=\"https://arxiv.org/pdf/2201.00057\" target=\"_blank\">OPTIMAL REPRESENTATIONS FOR COVARIATE SHIFT</a>.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F13324082%2Ffd8138c0e4a2c94fa71db54759b3c229%2FScreen%20Shot%202024-06-11%20at%2012.35.54%20PM.png?generation=1718080587087484&amp;alt=media\" alt=\"model\"><br>\nHonestly, I don't understand the math behind that(read the paper if you are interest). Practically, it use data domains as labels (in our case, train data: 0, unlabeled_soundscapes: 1), and bottleneck_layer's output as features to do SupCon learning. The total loss is traditional BCE loss + weight*SupCon loss. The weight and the temperature in SupCon loss is tuned base on public leaderboard.</p>\n<table>\n<thead>\n<tr>\n<th>weight</th>\n<th>Public</th>\n<th>Private</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>1e-2</td>\n<td>0.56</td>\n<td>0.57</td>\n</tr>\n<tr>\n<td>1e-3</td>\n<td>0.66</td>\n<td>0.63</td>\n</tr>\n<tr>\n<td>1e-4</td>\n<td>0.68</td>\n<td>0.66</td>\n</tr>\n<tr>\n<td>1e-5</td>\n<td>0.65</td>\n<td>0.64</td>\n</tr>\n</tbody>\n</table>\n<table>\n<thead>\n<tr>\n<th>temperature</th>\n<th>Public</th>\n<th>Private</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0.07</td>\n<td>0.62</td>\n<td>0.62</td>\n</tr>\n<tr>\n<td>0.09</td>\n<td>0.65</td>\n<td>0.65</td>\n</tr>\n<tr>\n<td>0.10</td>\n<td>0.68</td>\n<td>0.66</td>\n</tr>\n<tr>\n<td>0.11</td>\n<td>0.65</td>\n<td>0.66</td>\n</tr>\n</tbody>\n</table>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F13324082%2F99ef98a7c009cd712b421d73e84ed4d2%2FScreen%20Shot%202024-06-11%20at%201.00.20%20PM.png?generation=1718082202941471&amp;alt=media\" alt=\"supcon loss\"></p>\n<p>At the early stage, BCE loss dominates, then SupCon loss starts to kick in.</p>\n<h2><strong>Post-process trick</strong></h2>\n<p>Average prediction with previous and next window. Improvement depends(0.01~0.02).</p>\n<h2><strong>Model ensemble and selection</strong></h2>\n<p>My best ensemble is actually two eca_nfnet_l0s with different data split and SupCon loss weight, it scores 0.68 in private, but I didn't select it.<br>\nI further ensemble these two with a tf_efficientnetv2_b0(public 0.65, private 0.62, with time &amp; freq mask) and rexnet_100(public 0.63, private 0.64). Don't have enough time to finetune these two, maybe they can improve.</p>\n<h2><strong>Things didn't work(overfitting mostly)</strong></h2>\n<p>pretrain (on 22&amp;23 data, maybe I should try pretrain on unlabeled_soundscapes)<br>\nknowledge distillation (only try train data, maybe I should try knowledge distillation on unlabeled_soundscapes)<br>\nEntmin (definitely useless)<br>\nDANN (no improvement)<br>\nfocal loss (it actually works solely, but no improvement with CAD)</p>\n<h2><strong>Final thought</strong></h2>\n<p>This competition is very hard. On domain adaptation/generalization leaderboard Domainbed, simple baseline is hard to beat. In practice, no current methods for learning representation uniformly outperform the baseline. I think that's reason for the big shakeup, and why public notebooks outperform carefully trained models.</p>",
  "messages": [
    {
      "id": "2865837",
      "postDate": "06/11/2024 03:18:19",
      "content": "<p>Share some personal experiences in this competition.</p>\n<h2><strong>CV strategy</strong></h2>\n<p>No CV, only public leaderboard. It just works.<br>\nAt early stage, I tried to build CV and even wasted lots of submissions to do leaderboard probing. Then I realized that CV is just waste of time. I just assume unlabeled soundscapes, public test set and private test set share similar data distribution later on.</p>\n<h2><strong>Data Augmentation</strong></h2>\n<p>Only cutmix and mixup. Only one model with time &amp; freq mask and it overfits.<br>\nTraditional data augmentation tends to overfit the train data, because it teaches the model to output the same representation for data from single audio clip (single domain knowledge). While things like mixup have some cross domain knowledge, because data comes from different audio clips. (not really sure, just guessing)<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F13324082%2F5bda7d799dd9fc33b17de71030a35fbf%2FScreen%20Shot%202024-06-11%20at%2010.09.53%20AM.png?generation=1718071832807423&amp;alt=media\" alt=\"nice pic from OPTIMAL REPRESENTATIONS FOR COVARIATE SHIFT\"></p>\n<h2><strong>Some Model Design</strong></h2>\n<p>Almost last year's model. Add a bottleneck layer to let the representation more compact.<br>\n<code>self.bottleneck_layer = nn.Sequential(\n            nn.Linear(backbone_out, cfg.bottleneck_dim),\n            nn.BatchNorm1d(cfg.bottleneck_dim),\n            nn.ReLU())</code><br>\nbottleneck_dim=1024(actually 512 is better). This can improve score a little(less than 0.01).</p>\n<p>With eca_nfnet_l0 as backbone and above tips, my baseline public score is 0.65(close to 0.66).</p>\n<h2><strong>Contrastive Adversarial Domain (CAD) bottleneck</strong></h2>\n<p>from the paper <a href=\"https://arxiv.org/pdf/2201.00057\" target=\"_blank\">OPTIMAL REPRESENTATIONS FOR COVARIATE SHIFT</a>.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F13324082%2Ffd8138c0e4a2c94fa71db54759b3c229%2FScreen%20Shot%202024-06-11%20at%2012.35.54%20PM.png?generation=1718080587087484&amp;alt=media\" alt=\"model\"><br>\nHonestly, I don't understand the math behind that(read the paper if you are interest). Practically, it use data domains as labels (in our case, train data: 0, unlabeled_soundscapes: 1), and bottleneck_layer's output as features to do SupCon learning. The total loss is traditional BCE loss + weight*SupCon loss. The weight and the temperature in SupCon loss is tuned base on public leaderboard.</p>\n<table>\n<thead>\n<tr>\n<th>weight</th>\n<th>Public</th>\n<th>Private</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>1e-2</td>\n<td>0.56</td>\n<td>0.57</td>\n</tr>\n<tr>\n<td>1e-3</td>\n<td>0.66</td>\n<td>0.63</td>\n</tr>\n<tr>\n<td>1e-4</td>\n<td>0.68</td>\n<td>0.66</td>\n</tr>\n<tr>\n<td>1e-5</td>\n<td>0.65</td>\n<td>0.64</td>\n</tr>\n</tbody>\n</table>\n<table>\n<thead>\n<tr>\n<th>temperature</th>\n<th>Public</th>\n<th>Private</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0.07</td>\n<td>0.62</td>\n<td>0.62</td>\n</tr>\n<tr>\n<td>0.09</td>\n<td>0.65</td>\n<td>0.65</td>\n</tr>\n<tr>\n<td>0.10</td>\n<td>0.68</td>\n<td>0.66</td>\n</tr>\n<tr>\n<td>0.11</td>\n<td>0.65</td>\n<td>0.66</td>\n</tr>\n</tbody>\n</table>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F13324082%2F99ef98a7c009cd712b421d73e84ed4d2%2FScreen%20Shot%202024-06-11%20at%201.00.20%20PM.png?generation=1718082202941471&amp;alt=media\" alt=\"supcon loss\"></p>\n<p>At the early stage, BCE loss dominates, then SupCon loss starts to kick in.</p>\n<h2><strong>Post-process trick</strong></h2>\n<p>Average prediction with previous and next window. Improvement depends(0.01~0.02).</p>\n<h2><strong>Model ensemble and selection</strong></h2>\n<p>My best ensemble is actually two eca_nfnet_l0s with different data split and SupCon loss weight, it scores 0.68 in private, but I didn't select it.<br>\nI further ensemble these two with a tf_efficientnetv2_b0(public 0.65, private 0.62, with time &amp; freq mask) and rexnet_100(public 0.63, private 0.64). Don't have enough time to finetune these two, maybe they can improve.</p>\n<h2><strong>Things didn't work(overfitting mostly)</strong></h2>\n<p>pretrain (on 22&amp;23 data, maybe I should try pretrain on unlabeled_soundscapes)<br>\nknowledge distillation (only try train data, maybe I should try knowledge distillation on unlabeled_soundscapes)<br>\nEntmin (definitely useless)<br>\nDANN (no improvement)<br>\nfocal loss (it actually works solely, but no improvement with CAD)</p>\n<h2><strong>Final thought</strong></h2>\n<p>This competition is very hard. On domain adaptation/generalization leaderboard Domainbed, simple baseline is hard to beat. In practice, no current methods for learning representation uniformly outperform the baseline. I think that's reason for the big shakeup, and why public notebooks outperform carefully trained models.</p>",
      "rawMarkdown": "Share some personal experiences in this competition.\n\n## **CV strategy** ##\nNo CV, only public leaderboard. It just works.\nAt early stage, I tried to build CV and even wasted lots of submissions to do leaderboard probing. Then I realized that CV is just waste of time. I just assume unlabeled soundscapes, public test set and private test set share similar data distribution later on.\n\n## **Data Augmentation** ##\nOnly cutmix and mixup. Only one model with time & freq mask and it overfits.\nTraditional data augmentation tends to overfit the train data, because it teaches the model to output the same representation for data from single audio clip (single domain knowledge). While things like mixup have some cross domain knowledge, because data comes from different audio clips. (not really sure, just guessing)\n![nice pic from OPTIMAL REPRESENTATIONS FOR COVARIATE SHIFT](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F13324082%2F5bda7d799dd9fc33b17de71030a35fbf%2FScreen%20Shot%202024-06-11%20at%2010.09.53%20AM.png?generation=1718071832807423&alt=media ''nice pic from OPTIMAL REPRESENTATIONS FOR COVARIATE SHIFT'')\n\n## **Some Model Design** ##\nAlmost last year's model. Add a bottleneck layer to let the representation more compact.\n`self.bottleneck_layer = nn.Sequential(\n            nn.Linear(backbone_out, cfg.bottleneck_dim),\n            nn.BatchNorm1d(cfg.bottleneck_dim),\n            nn.ReLU())`\nbottleneck_dim=1024(actually 512 is better). This can improve score a little(less than 0.01).\n\nWith eca_nfnet_l0 as backbone and above tips, my baseline public score is 0.65(close to 0.66).\n\n## **Contrastive Adversarial Domain (CAD) bottleneck** ##\nfrom the paper [OPTIMAL REPRESENTATIONS FOR COVARIATE SHIFT](https://arxiv.org/pdf/2201.00057).\n![model](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F13324082%2Ffd8138c0e4a2c94fa71db54759b3c229%2FScreen%20Shot%202024-06-11%20at%2012.35.54%20PM.png?generation=1718080587087484&alt=media ''model' structure'')\nHonestly, I don't understand the math behind that(read the paper if you are interest). Practically, it use data domains as labels (in our case, train data: 0, unlabeled_soundscapes: 1), and bottleneck_layer's output as features to do SupCon learning. The total loss is traditional BCE loss + weight*SupCon loss. The weight and the temperature in SupCon loss is tuned base on public leaderboard.\n\n| weight         | Public           | Private |\n| ------------- |:-------------:| -----:|\n| 1e-2 | 0.56 | 0.57 |\n| 1e-3 | 0.66 | 0.63 |\n| 1e-4 | 0.68 | 0.66 |\n| 1e-5 | 0.65 | 0.64 |\n\n\n| temperature | Public           | Private |\n| ------------- |:-------------:| -----:|\n| 0.07 | 0.62 | 0.62 |\n| 0.09 | 0.65 | 0.65 |\n| 0.10 | 0.68 | 0.66 |\n| 0.11 | 0.65 | 0.66 |\n\n![supcon loss](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F13324082%2F99ef98a7c009cd712b421d73e84ed4d2%2FScreen%20Shot%202024-06-11%20at%201.00.20%20PM.png?generation=1718082202941471&alt=media ''train SupCon loss'')\n\nAt the early stage, BCE loss dominates, then SupCon loss starts to kick in.\n\n## **Post-process trick** ##\nAverage prediction with previous and next window. Improvement depends(0.01~0.02).\n\n## **Model ensemble and selection** ##\nMy best ensemble is actually two eca_nfnet_l0s with different data split and SupCon loss weight, it scores 0.68 in private, but I didn't select it.\nI further ensemble these two with a tf_efficientnetv2_b0(public 0.65, private 0.62, with time & freq mask) and rexnet_100(public 0.63, private 0.64). Don't have enough time to finetune these two, maybe they can improve.\n\n## **Things didn't work(overfitting mostly)** ##\npretrain (on 22&23 data, maybe I should try pretrain on unlabeled_soundscapes)\nknowledge distillation (only try train data, maybe I should try knowledge distillation on unlabeled_soundscapes)\nEntmin (definitely useless)\nDANN (no improvement)\nfocal loss (it actually works solely, but no improvement with CAD)\n\n## **Final thought** ##\nThis competition is very hard. On domain adaptation/generalization leaderboard Domainbed, simple baseline is hard to beat. In practice, no current methods for learning representation uniformly outperform the baseline. I think that's reason for the big shakeup, and why public notebooks outperform carefully trained models.",
      "votes": null
    },
    {
      "id": "2865939",
      "postDate": "06/11/2024 04:35:16",
      "content": "<p>Congratulations on achieving 9th place in this competition. Thanks for sharing your insights. </p>",
      "rawMarkdown": "Congratulations on achieving 9th place in this competition. Thanks for sharing your insights.",
      "votes": null
    },
    {
      "id": "2866257",
      "postDate": "06/11/2024 08:22:28",
      "content": "<p>Congratsss!!! </p>",
      "rawMarkdown": "Congratsss!!!",
      "votes": null
    },
    {
      "id": "2866623",
      "postDate": "06/11/2024 12:14:02",
      "content": "<p>Congratulations on your achievement! Thank you for your detailed explanation with diagrams.</p>\n<p>This was my first time participating in a Kaggle competition, and through the experience, I developed a strong interest in domain adaptation. You mentioned that you successfully used CAD to address domain shift, while DANN was not as effective. How do you approach the selection of methods for domain adaptation?</p>\n<p>I have tried three domain adaptation methods: Cycle GAN, DANN, and M2D. The top private scores in my submissions were achieved using Cycle GAN, suggesting some level of effectiveness, but DANN did not perform well.</p>\n<p>Despite the significant shake-up at the end, I am determined to learn and improve. Thank you for your time and insights.</p>",
      "rawMarkdown": "Congratulations on your achievement! Thank you for your detailed explanation with diagrams.\n\nThis was my first time participating in a Kaggle competition, and through the experience, I developed a strong interest in domain adaptation. You mentioned that you successfully used CAD to address domain shift, while DANN was not as effective. How do you approach the selection of methods for domain adaptation?\n\nI have tried three domain adaptation methods: Cycle GAN, DANN, and M2D. The top private scores in my submissions were achieved using Cycle GAN, suggesting some level of effectiveness, but DANN did not perform well.\n\nDespite the significant shake-up at the end, I am determined to learn and improve. Thank you for your time and insights.",
      "votes": null
    },
    {
      "id": "2866648",
      "postDate": "06/11/2024 12:38:35",
      "content": "<p>I can recommend a <a href=\"https://www.cs.purdue.edu/homes/ribeirob/courses/Spring2023/lectures/12CovShiftandOOD/domain_covariate_shift.html\" target=\"_blank\">reading material</a> which leads me to the CAD paper. As for method selection, just trial and error. Like I said, no current methods uniformly outperform the baseline. The downside of DANN is, ''(i) it<br>\nmaximizes an upper bound on the desired term;<br>\n(ii) it requires adversarial training, which is challenging in practice'', quoted from CAD paper. You should also try <a href=\"https://github.com/facebookresearch/DomainBed\" target=\"_blank\">DomainBed</a> which contain lots of domain adaptation tricks. </p>",
      "rawMarkdown": "I can recommend a [reading material](https://www.cs.purdue.edu/homes/ribeirob/courses/Spring2023/lectures/12CovShiftandOOD/domain_covariate_shift.html) which leads me to the CAD paper. As for method selection, just trial and error. Like I said, no current methods uniformly outperform the baseline. The downside of DANN is, ''(i) it\nmaximizes an upper bound on the desired term;\n(ii) it requires adversarial training, which is challenging in practice'', quoted from CAD paper. You should also try [DomainBed](https://github.com/facebookresearch/DomainBed) which contain lots of domain adaptation tricks.",
      "votes": null
    },
    {
      "id": "2866667",
      "postDate": "06/11/2024 12:50:23",
      "content": "<p>Thank you for your response and the recommendation of reading material that led you to the CAD paper. I will definitely look into it.</p>\n<p>As for the method selection, I appreciate your insight about trial and error. It's reassuring to know that even experienced practitioners face challenges with adversarial training. I found the quote from the CAD paper about the downsides of DANN particularly enlightening.</p>\n<p>I will also explore DomainBed as you suggested. It's good to know that it contains various domain adaptation tricks that could be useful for my future projects.</p>\n<p>Thank you once again for your guidance and recommendations. Your insights are incredibly valuable to someone new to these methods.</p>",
      "rawMarkdown": "Thank you for your response and the recommendation of reading material that led you to the CAD paper. I will definitely look into it.\n\nAs for the method selection, I appreciate your insight about trial and error. It's reassuring to know that even experienced practitioners face challenges with adversarial training. I found the quote from the CAD paper about the downsides of DANN particularly enlightening.\n\nI will also explore DomainBed as you suggested. It's good to know that it contains various domain adaptation tricks that could be useful for my future projects.\n\nThank you once again for your guidance and recommendations. Your insights are incredibly valuable to someone new to these methods.",
      "votes": null
    },
    {
      "id": "2878303",
      "postDate": "06/18/2024 21:09:33",
      "content": "<p>Cool approach! Really nice to see someone try a domain adaptation method in the competition context.</p>",
      "rawMarkdown": "Cool approach! Really nice to see someone try a domain adaptation method in the competition context.",
      "votes": null
    },
    {
      "id": "2878542",
      "postDate": "06/19/2024 03:41:27",
      "content": "<p>Thanks! Glad you like it.</p>",
      "rawMarkdown": "Thanks! Glad you like it.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2865939,
      "author_name": "crsuthikshnkumar",
      "author_url": "",
      "post_date": "06/11/2024 04:35:16",
      "content": "<p>Congratulations on achieving 9th place in this competition. Thanks for sharing your insights. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2866257,
      "author_name": "aadityaporwal",
      "author_url": "",
      "post_date": "06/11/2024 08:22:28",
      "content": "<p>Congratsss!!! </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2866623,
      "author_name": "eiichiokumura",
      "author_url": "",
      "post_date": "06/11/2024 12:14:02",
      "content": "<p>Congratulations on your achievement! Thank you for your detailed explanation with diagrams.</p>\n<p>This was my first time participating in a Kaggle competition, and through the experience, I developed a strong interest in domain adaptation. You mentioned that you successfully used CAD to address domain shift, while DANN was not as effective. How do you approach the selection of methods for domain adaptation?</p>\n<p>I have tried three domain adaptation methods: Cycle GAN, DANN, and M2D. The top private scores in my submissions were achieved using Cycle GAN, suggesting some level of effectiveness, but DANN did not perform well.</p>\n<p>Despite the significant shake-up at the end, I am determined to learn and improve. Thank you for your time and insights.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2866648,
          "author_name": "aphysict",
          "author_url": "",
          "post_date": "06/11/2024 12:38:35",
          "content": "<p>I can recommend a <a href=\"https://www.cs.purdue.edu/homes/ribeirob/courses/Spring2023/lectures/12CovShiftandOOD/domain_covariate_shift.html\" target=\"_blank\">reading material</a> which leads me to the CAD paper. As for method selection, just trial and error. Like I said, no current methods uniformly outperform the baseline. The downside of DANN is, ''(i) it<br>\nmaximizes an upper bound on the desired term;<br>\n(ii) it requires adversarial training, which is challenging in practice'', quoted from CAD paper. You should also try <a href=\"https://github.com/facebookresearch/DomainBed\" target=\"_blank\">DomainBed</a> which contain lots of domain adaptation tricks. </p>",
          "votes": null,
          "replies": [
            {
              "id": 2866667,
              "author_name": "eiichiokumura",
              "author_url": "",
              "post_date": "06/11/2024 12:50:23",
              "content": "<p>Thank you for your response and the recommendation of reading material that led you to the CAD paper. I will definitely look into it.</p>\n<p>As for the method selection, I appreciate your insight about trial and error. It's reassuring to know that even experienced practitioners face challenges with adversarial training. I found the quote from the CAD paper about the downsides of DANN particularly enlightening.</p>\n<p>I will also explore DomainBed as you suggested. It's good to know that it contains various domain adaptation tricks that could be useful for my future projects.</p>\n<p>Thank you once again for your guidance and recommendations. Your insights are incredibly valuable to someone new to these methods.</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2878303,
      "author_name": "tomdenton",
      "author_url": "",
      "post_date": "06/18/2024 21:09:33",
      "content": "<p>Cool approach! Really nice to see someone try a domain adaptation method in the competition context.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2878542,
          "author_name": "aphysict",
          "author_url": "",
          "post_date": "06/19/2024 03:41:27",
          "content": "<p>Thanks! Glad you like it.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2865837": "Share some personal experiences in this competition.\n\n## **CV strategy** ##\nNo CV, only public leaderboard. It just works.\nAt early stage, I tried to build CV and even wasted lots of submissions to do leaderboard probing. Then I realized that CV is just waste of time. I just assume unlabeled soundscapes, public test set and private test set share similar data distribution later on.\n\n## **Data Augmentation** ##\nOnly cutmix and mixup. Only one model with time & freq mask and it overfits.\nTraditional data augmentation tends to overfit the train data, because it teaches the model to output the same representation for data from single audio clip (single domain knowledge). While things like mixup have some cross domain knowledge, because data comes from different audio clips. (not really sure, just guessing)\n![nice pic from OPTIMAL REPRESENTATIONS FOR COVARIATE SHIFT](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F13324082%2F5bda7d799dd9fc33b17de71030a35fbf%2FScreen%20Shot%202024-06-11%20at%2010.09.53%20AM.png?generation=1718071832807423&alt=media ''nice pic from OPTIMAL REPRESENTATIONS FOR COVARIATE SHIFT'')\n\n## **Some Model Design** ##\nAlmost last year's model. Add a bottleneck layer to let the representation more compact.\n`self.bottleneck_layer = nn.Sequential(\n            nn.Linear(backbone_out, cfg.bottleneck_dim),\n            nn.BatchNorm1d(cfg.bottleneck_dim),\n            nn.ReLU())`\nbottleneck_dim=1024(actually 512 is better). This can improve score a little(less than 0.01).\n\nWith eca_nfnet_l0 as backbone and above tips, my baseline public score is 0.65(close to 0.66).\n\n## **Contrastive Adversarial Domain (CAD) bottleneck** ##\nfrom the paper [OPTIMAL REPRESENTATIONS FOR COVARIATE SHIFT](https://arxiv.org/pdf/2201.00057).\n![model](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F13324082%2Ffd8138c0e4a2c94fa71db54759b3c229%2FScreen%20Shot%202024-06-11%20at%2012.35.54%20PM.png?generation=1718080587087484&alt=media ''model' structure'')\nHonestly, I don't understand the math behind that(read the paper if you are interest). Practically, it use data domains as labels (in our case, train data: 0, unlabeled_soundscapes: 1), and bottleneck_layer's output as features to do SupCon learning. The total loss is traditional BCE loss + weight*SupCon loss. The weight and the temperature in SupCon loss is tuned base on public leaderboard.\n\n| weight         | Public           | Private |\n| ------------- |:-------------:| -----:|\n| 1e-2 | 0.56 | 0.57 |\n| 1e-3 | 0.66 | 0.63 |\n| 1e-4 | 0.68 | 0.66 |\n| 1e-5 | 0.65 | 0.64 |\n\n\n| temperature | Public           | Private |\n| ------------- |:-------------:| -----:|\n| 0.07 | 0.62 | 0.62 |\n| 0.09 | 0.65 | 0.65 |\n| 0.10 | 0.68 | 0.66 |\n| 0.11 | 0.65 | 0.66 |\n\n![supcon loss](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F13324082%2F99ef98a7c009cd712b421d73e84ed4d2%2FScreen%20Shot%202024-06-11%20at%201.00.20%20PM.png?generation=1718082202941471&alt=media ''train SupCon loss'')\n\nAt the early stage, BCE loss dominates, then SupCon loss starts to kick in.\n\n## **Post-process trick** ##\nAverage prediction with previous and next window. Improvement depends(0.01~0.02).\n\n## **Model ensemble and selection** ##\nMy best ensemble is actually two eca_nfnet_l0s with different data split and SupCon loss weight, it scores 0.68 in private, but I didn't select it.\nI further ensemble these two with a tf_efficientnetv2_b0(public 0.65, private 0.62, with time & freq mask) and rexnet_100(public 0.63, private 0.64). Don't have enough time to finetune these two, maybe they can improve.\n\n## **Things didn't work(overfitting mostly)** ##\npretrain (on 22&23 data, maybe I should try pretrain on unlabeled_soundscapes)\nknowledge distillation (only try train data, maybe I should try knowledge distillation on unlabeled_soundscapes)\nEntmin (definitely useless)\nDANN (no improvement)\nfocal loss (it actually works solely, but no improvement with CAD)\n\n## **Final thought** ##\nThis competition is very hard. On domain adaptation/generalization leaderboard Domainbed, simple baseline is hard to beat. In practice, no current methods for learning representation uniformly outperform the baseline. I think that's reason for the big shakeup, and why public notebooks outperform carefully trained models.",
    "2865939": "Congratulations on achieving 9th place in this competition. Thanks for sharing your insights.",
    "2866257": "Congratsss!!!",
    "2866623": "Congratulations on your achievement! Thank you for your detailed explanation with diagrams.\n\nThis was my first time participating in a Kaggle competition, and through the experience, I developed a strong interest in domain adaptation. You mentioned that you successfully used CAD to address domain shift, while DANN was not as effective. How do you approach the selection of methods for domain adaptation?\n\nI have tried three domain adaptation methods: Cycle GAN, DANN, and M2D. The top private scores in my submissions were achieved using Cycle GAN, suggesting some level of effectiveness, but DANN did not perform well.\n\nDespite the significant shake-up at the end, I am determined to learn and improve. Thank you for your time and insights.",
    "2866648": "I can recommend a [reading material](https://www.cs.purdue.edu/homes/ribeirob/courses/Spring2023/lectures/12CovShiftandOOD/domain_covariate_shift.html) which leads me to the CAD paper. As for method selection, just trial and error. Like I said, no current methods uniformly outperform the baseline. The downside of DANN is, ''(i) it\nmaximizes an upper bound on the desired term;\n(ii) it requires adversarial training, which is challenging in practice'', quoted from CAD paper. You should also try [DomainBed](https://github.com/facebookresearch/DomainBed) which contain lots of domain adaptation tricks.",
    "2866667": "Thank you for your response and the recommendation of reading material that led you to the CAD paper. I will definitely look into it.\n\nAs for the method selection, I appreciate your insight about trial and error. It's reassuring to know that even experienced practitioners face challenges with adversarial training. I found the quote from the CAD paper about the downsides of DANN particularly enlightening.\n\nI will also explore DomainBed as you suggested. It's good to know that it contains various domain adaptation tricks that could be useful for my future projects.\n\nThank you once again for your guidance and recommendations. Your insights are incredibly valuable to someone new to these methods.",
    "2878303": "Cool approach! Really nice to see someone try a domain adaptation method in the competition context.",
    "2878542": "Thanks! Glad you like it."
  },
  "source": "meta"
}