{
  "id": 266397,
  "title": "2nd Place Solution",
  "url": "/competitions/seti-breakthrough-listen/discussion/266397",
  "author_name": "hirune924",
  "post_date": "2021-08-19T01:47:38.143000",
  "votes": 74,
  "comment_count": 38,
  "views": 0,
  "content": "<p>Thanks to Kaggle and Berkeley SETI Research Center for this competition. Finding the aliens was very exciting. I tried many things to close the gap between cv and lb, but few succeeded, and as a result, my solution is very simple.</p>\n<h2>Abstract</h2>\n<p>My solution is a two-step process. The first stage is training with train data only, and the next stage is training with test data and pseudo labels. In the first stage, I use only vflip, cutout, and mixup to prevent the model from being confused by losing the signal due to unexpected augmentation. In addition, StochasticDepth and Dropout are strongly used to prevent overtraining. The mixup uses logical OR instead of alpha blending when mixing targets. This allows creating a model that responds strongly to weak signals. At this stage, we achieved PublicLB 0.800 by training efficientnetB5.</p>\n<p>In the second stage, we finetune the first stage model using pseudo-labels created by using the pre-trained model from the first stage. In this case, we refer to NoisyStudent and use Stochastic Depth and Dropout more strongly to prevent overfitting to the pseudo-label. By repeating this process several times, the PublicLB 0.813 is reached. Finally, I simply averaged the five submissions with the highest LB scores.</p>\n<h2>1st Stage Details</h2>\n<p>You can see a sample of the code here.<br>\n<a href=\"https://www.kaggle.com/hirune924/2ndplace-solution\" target=\"_blank\">https://www.kaggle.com/hirune924/2ndplace-solution</a></p>\n<ul>\n<li>The input images are merged in the time direction and resized to a size of 512x512.</li>\n<li>The mixup target can be mixed by using the following expression to express a logical OR, which also supports soft targets when using pseudo labels.</li>\n</ul>\n<pre><code>y = y + y[index] - (y * y[index])\n</code></pre>\n<ul>\n<li>TTA with vflip is used for inference</li>\n</ul>\n<p>Sometimes, the loss became nan or the validation auc became 0.5 and the training did not progress. In this case, I solved the problem by training only the head and encoder BNs for one epoch first.</p>\n<h2>2nd Stage Details</h2>\n<p>The code is almost the same as the first step, adding pseudo-labels and making Stochastic Depth and Dropout even stronger for training.</p>\n<h2>Didn't work</h2>\n<ul>\n<li>Unsupervised Domain Adaptation by finetune only BN</li>\n<li>Add artificial signal, chi2noise, referring to setigen.</li>\n<li>input stem stride=1</li>\n<li>using old data</li>\n<li>Anomaly detection based metric learning</li>\n<li>Semi Supervised Learning</li>\n<li>many other things…</li>\n</ul>",
  "messages": [
    {
      "id": 1480362,
      "postDate": "2021-08-19T01:47:38.143Z",
      "content": "<p>Thanks to Kaggle and Berkeley SETI Research Center for this competition. Finding the aliens was very exciting. I tried many things to close the gap between cv and lb, but few succeeded, and as a result, my solution is very simple.</p>\n<h2>Abstract</h2>\n<p>My solution is a two-step process. The first stage is training with train data only, and the next stage is training with test data and pseudo labels. In the first stage, I use only vflip, cutout, and mixup to prevent the model from being confused by losing the signal due to unexpected augmentation. In addition, StochasticDepth and Dropout are strongly used to prevent overtraining. The mixup uses logical OR instead of alpha blending when mixing targets. This allows creating a model that responds strongly to weak signals. At this stage, we achieved PublicLB 0.800 by training efficientnetB5.</p>\n<p>In the second stage, we finetune the first stage model using pseudo-labels created by using the pre-trained model from the first stage. In this case, we refer to NoisyStudent and use Stochastic Depth and Dropout more strongly to prevent overfitting to the pseudo-label. By repeating this process several times, the PublicLB 0.813 is reached. Finally, I simply averaged the five submissions with the highest LB scores.</p>\n<h2>1st Stage Details</h2>\n<p>You can see a sample of the code here.<br>\n<a href=\"https://www.kaggle.com/hirune924/2ndplace-solution\" target=\"_blank\">https://www.kaggle.com/hirune924/2ndplace-solution</a></p>\n<ul>\n<li>The input images are merged in the time direction and resized to a size of 512x512.</li>\n<li>The mixup target can be mixed by using the following expression to express a logical OR, which also supports soft targets when using pseudo labels.</li>\n</ul>\n<pre><code>y = y + y[index] - (y * y[index])\n</code></pre>\n<ul>\n<li>TTA with vflip is used for inference</li>\n</ul>\n<p>Sometimes, the loss became nan or the validation auc became 0.5 and the training did not progress. In this case, I solved the problem by training only the head and encoder BNs for one epoch first.</p>\n<h2>2nd Stage Details</h2>\n<p>The code is almost the same as the first step, adding pseudo-labels and making Stochastic Depth and Dropout even stronger for training.</p>\n<h2>Didn't work</h2>\n<ul>\n<li>Unsupervised Domain Adaptation by finetune only BN</li>\n<li>Add artificial signal, chi2noise, referring to setigen.</li>\n<li>input stem stride=1</li>\n<li>using old data</li>\n<li>Anomaly detection based metric learning</li>\n<li>Semi Supervised Learning</li>\n<li>many other things…</li>\n</ul>",
      "rawMarkdown": "Thanks to Kaggle and Berkeley SETI Research Center for this competition. Finding the aliens was very exciting. I tried many things to close the gap between cv and lb, but few succeeded, and as a result, my solution is very simple.\n\n## Abstract\nMy solution is a two-step process. The first stage is training with train data only, and the next stage is training with test data and pseudo labels. In the first stage, I use only vflip, cutout, and mixup to prevent the model from being confused by losing the signal due to unexpected augmentation. In addition, StochasticDepth and Dropout are strongly used to prevent overtraining. The mixup uses logical OR instead of alpha blending when mixing targets. This allows creating a model that responds strongly to weak signals. At this stage, we achieved PublicLB 0.800 by training efficientnetB5.\n\nIn the second stage, we finetune the first stage model using pseudo-labels created by using the pre-trained model from the first stage. In this case, we refer to NoisyStudent and use Stochastic Depth and Dropout more strongly to prevent overfitting to the pseudo-label. By repeating this process several times, the PublicLB 0.813 is reached. Finally, I simply averaged the five submissions with the highest LB scores.\n\n## 1st Stage Details\nYou can see a sample of the code here.\nhttps://www.kaggle.com/hirune924/2ndplace-solution\n* The input images are merged in the time direction and resized to a size of 512x512.\n* The mixup target can be mixed by using the following expression to express a logical OR, which also supports soft targets when using pseudo labels.\n```\ny = y + y[index] - (y * y[index])\n```\n* TTA with vflip is used for inference\n\nSometimes, the loss became nan or the validation auc became 0.5 and the training did not progress. In this case, I solved the problem by training only the head and encoder BNs for one epoch first.\n\n## 2nd Stage Details\nThe code is almost the same as the first step, adding pseudo-labels and making Stochastic Depth and Dropout even stronger for training.\n\n## Didn't work\n* Unsupervised Domain Adaptation by finetune only BN\n* Add artificial signal, chi2noise, referring to setigen.\n* input stem stride=1\n* using old data\n* Anomaly detection based metric learning\n* Semi Supervised Learning\n* many other things...\n\n",
      "votes": 74
    },
    {
      "id": 1480685,
      "postDate": "2021-08-19T06:17:21.457Z",
      "content": "<p>Congrats, very impressive results. When we saw you climbing one point every day, we were quite convinced that you are pseudo tagging :)</p>\n<p>The important part here seems to do heavy regularization when incorporating the pseudo tags, so that the model does not overfit too heavily on the pseudo tags coming from images of a different distribution, we shortly discuss this also in our solution post. I believe that the best way to incorporate pseudo tags, in theory, would be to only add target=1 labels, because, at least from CV, the models make less mistakes when they predict and find the signal, compared to when they don't. However, this does not work here, because then the model would just learn that target=1 comes always from test (due to distribution shift). There might be some clever way to do this, by studying more the research on accounting for distribution shift (e.g., additional losses that try to prevent overfitting on the distribution). Unfortunately, didnt have time to look into it due to our other findings.</p>\n<p>Again: very impressive job and congrats on solo gold!</p>",
      "rawMarkdown": "Congrats, very impressive results. When we saw you climbing one point every day, we were quite convinced that you are pseudo tagging :)\n\nThe important part here seems to do heavy regularization when incorporating the pseudo tags, so that the model does not overfit too heavily on the pseudo tags coming from images of a different distribution, we shortly discuss this also in our solution post. I believe that the best way to incorporate pseudo tags, in theory, would be to only add target=1 labels, because, at least from CV, the models make less mistakes when they predict and find the signal, compared to when they don't. However, this does not work here, because then the model would just learn that target=1 comes always from test (due to distribution shift). There might be some clever way to do this, by studying more the research on accounting for distribution shift (e.g., additional losses that try to prevent overfitting on the distribution). Unfortunately, didnt have time to look into it due to our other findings.\n\nAgain: very impressive job and congrats on solo gold!",
      "votes": 11,
      "replies": [
        {
          "id": 1480741,
          "postDate": "2021-08-19T06:48:19.727Z",
          "content": "<p>Your team's solution is also a great job. Congratulations on the 1st place.</p>\n<p>Actually, I thought that using all the pseudo-labels instead of only target=1 is dangerous because it may overfit to the pseudo-labels.<br>\nTherefore, we tried to use only those pseudo-labels that showed extreme target scores by threshold, but the Noisy Student style method using all pseudo-labels showed better and more stable results.<br>\nI would have liked to increase the model size according to the paper, but there was not enough time.</p>",
          "rawMarkdown": "Your team's solution is also a great job. Congratulations on the 1st place.\n\nActually, I thought that using all the pseudo-labels instead of only target=1 is dangerous because it may overfit to the pseudo-labels.\nTherefore, we tried to use only those pseudo-labels that showed extreme target scores by threshold, but the Noisy Student style method using all pseudo-labels showed better and more stable results.\nI would have liked to increase the model size according to the paper, but there was not enough time.",
          "votes": 5
        }
      ]
    },
    {
      "id": 1480428,
      "postDate": "2021-08-19T02:56:21.657Z",
      "content": "<p>Congratulations! your boost on 2nd stage is very inpressive!</p>\n<blockquote>\n  <p>y = y + y[index] - (y * y[index])</p>\n</blockquote>\n<p>This is completely the same as our mixup: <code>torch.stack([y, y[index]], 0).max(0).values</code></p>\n<p>But we didn't adopted PL or NS…</p>",
      "rawMarkdown": "Congratulations! your boost on 2nd stage is very inpressive!\n\n> y = y + y[index] - (y * y[index])\n\nThis is completely the same as our mixup: `torch.stack([y, y[index]], 0).max(0).values`\n\nBut we didn't adopted PL or NS...",
      "votes": 5,
      "replies": [
        {
          "id": 1480737,
          "postDate": "2021-08-19T06:45:27.797Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/haqishen\" target=\"_blank\">@haqishen</a> my intuition was that if we use mixup it will only distort the signal and will not be good . I am still very weak in CV can you please tell us the thought process(intuition) behind using mixup and the trick used above</p>",
          "rawMarkdown": "Hi @haqishen my intuition was that if we use mixup it will only distort the signal and will not be good . I am still very weak in CV can you please tell us the thought process(intuition) behind using mixup and the trick used above"
        },
        {
          "id": 1481066,
          "postDate": "2021-08-19T09:42:49.290Z",
          "content": "<p>Actually weaken the injected signal is what we want, because we need the model to gain the ability to detect weak signal as well. The modified mixup is base on this intuition.</p>",
          "rawMarkdown": "Actually weaken the injected signal is what we want, because we need the model to gain the ability to detect weak signal as well. The modified mixup is base on this intuition.",
          "votes": 4
        },
        {
          "id": 1481145,
          "postDate": "2021-08-19T10:42:21.307Z",
          "content": "<p>Thank you </p>",
          "rawMarkdown": "Thank you "
        }
      ]
    },
    {
      "id": 1480397,
      "postDate": "2021-08-19T02:26:27.783Z",
      "content": "<p>Congrats!  Your solution looks quite similar to ours, but you gain more from pseudo labeling test.  I wonder why. Interestingly I used a very similar trick for target in mixup:</p>\n<pre><code>            target = alpha * target + (1 - alpha) * target[perm]\n            target = torch.clamp(2*target, 0, 1)\n</code></pre>",
      "rawMarkdown": "Congrats!  Your solution looks quite similar to ours, but you gain more from pseudo labeling test.  I wonder why. Interestingly I used a very similar trick for target in mixup:\n\n                target = alpha * target + (1 - alpha) * target[perm]\n                target = torch.clamp(2*target, 0, 1)\n",
      "votes": 5,
      "replies": [
        {
          "id": 1480430,
          "postDate": "2021-08-19T02:56:38.207Z",
          "content": "<p>I'm honored to have arrived at the same trick as you !</p>",
          "rawMarkdown": "I'm honored to have arrived at the same trick as you !",
          "votes": 2
        },
        {
          "id": 1481175,
          "postDate": "2021-08-19T11:02:06.487Z",
          "content": "<p>I am the one who is honored given you did much better than us!</p>",
          "rawMarkdown": "I am the one who is honored given you did much better than us!",
          "votes": 1
        },
        {
          "id": 1482292,
          "postDate": "2021-08-20T02:23:21.267Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 1482342,
          "postDate": "2021-08-20T03:26:46.580Z",
          "content": "<p>let say alpha is 0.5 i.e. the new image is the mean of the two mixed images. If one of them is of target 1 and the other 0, the mean target is 0.5.  By multiplying by 2 it becomes 1=, which is what we want here.</p>",
          "rawMarkdown": "let say alpha is 0.5 i.e. the new image is the mean of the two mixed images. If one of them is of target 1 and the other 0, the mean target is 0.5.  By multiplying by 2 it becomes 1=, which is what we want here.",
          "votes": 1
        },
        {
          "id": 1482385,
          "postDate": "2021-08-20T04:00:37.020Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> , regarding how to mix the target value during Mixup, sometimes people use standard(weighted) way like original <a href=\"https://arxiv.org/pdf/1710.09412v2.pdf\" target=\"_blank\">paper</a>, sometimes use maximum(logical OR) way like you did, I'm wondering what is the benefit for each way? The latter way is better for datasets which have very few positive samples, or target singal is very weak like this competition?</p>",
          "rawMarkdown": "Hi @cpmpml , regarding how to mix the target value during Mixup, sometimes people use standard(weighted) way like original [paper](https://arxiv.org/pdf/1710.09412v2.pdf), sometimes use maximum(logical OR) way like you did, I'm wondering what is the benefit for each way? The latter way is better for datasets which have very few positive samples, or target singal is very weak like this competition?"
        },
        {
          "id": 1483260,
          "postDate": "2021-08-20T14:36:55.963Z",
          "content": "<p>Here we look for messages.  A message stays in the mixed image, and the hope is that this way model will learn to recognized weaker signals (weaker because they are blended with another image).</p>",
          "rawMarkdown": "Here we look for messages.  A message stays in the mixed image, and the hope is that this way model will learn to recognized weaker signals (weaker because they are blended with another image).\n\n",
          "votes": 1
        }
      ]
    },
    {
      "id": 1480373,
      "postDate": "2021-08-19T02:04:17.630Z",
      "content": "<p>Thanks for sharing!!　Congratulations on the solo gold medal!!<br>\nI also tried arcface(m=0.5) to cope with unseen cls, but it failed. it takes 2x times than normal training and score wasn't good(LB=0.776,B0, size=768<em>768</em>3 )</p>",
      "rawMarkdown": "Thanks for sharing!!　Congratulations on the solo gold medal!!\nI also tried arcface(m=0.5) to cope with unseen cls, but it failed. it takes 2x times than normal training and score wasn't good(LB=0.776,B0, size=768*768*3 )",
      "votes": 3,
      "replies": [
        {
          "id": 1480434,
          "postDate": "2021-08-19T02:59:06.740Z",
          "content": "<p>I tried L2 constrained softmax Loss, but the max I could get was about lb0.73.</p>",
          "rawMarkdown": "I tried L2 constrained softmax Loss, but the max I could get was about lb0.73.",
          "votes": 1
        }
      ]
    },
    {
      "id": 1495846,
      "postDate": "2021-08-29T20:40:49.727Z",
      "content": "<p>Great work! Congrats</p>",
      "rawMarkdown": "Great work! Congrats",
      "votes": 1
    },
    {
      "id": 1480504,
      "postDate": "2021-08-19T03:56:00.087Z",
      "content": "<p>Congratulations for the solo gold <a href=\"https://www.kaggle.com/hirune924\" target=\"_blank\">@hirune924</a>  🎉</p>\n<blockquote>\n  <p>the loss became nan or the validation auc became 0.5 and the training did not progress. In this case, I solved the problem by training only the head and encoder BNs for one epoch first.</p>\n</blockquote>\n<p>I also observed sometimes my model does not learn anything as validation AUC is alway ~0.5, especially when I add more augmentations like ShiftScaleRotate or RandomResizedCrop, finally I found switch to OneCycleLR scheduler would help a bit, I suspect it`s due to it will start training with a small LR(1e-6 eg.) to find corrent direction to converge. May I know what is your theory about this \"always 0.5 CV\" problem? And why training with only head and encoder BNs in 1st epoch would help? Thanks.</p>",
      "rawMarkdown": "Congratulations for the solo gold @hirune924  🎉\n\n> the loss became nan or the validation auc became 0.5 and the training did not progress. In this case, I solved the problem by training only the head and encoder BNs for one epoch first.\n\nI also observed sometimes my model does not learn anything as validation AUC is alway ~0.5, especially when I add more augmentations like ShiftScaleRotate or RandomResizedCrop, finally I found switch to OneCycleLR scheduler would help a bit, I suspect it`s due to it will start training with a small LR(1e-6 eg.) to find corrent direction to converge. May I know what is your theory about this \"always 0.5 CV\" problem? And why training with only head and encoder BNs in 1st epoch would help? Thanks.",
      "votes": 1,
      "replies": [
        {
          "id": 1480528,
          "postDate": "2021-08-19T04:16:26.527Z",
          "content": "<p>Typically, we use imagenet pre-trained models for the encoder and untrained Linear for the classifier.　In the early stage of training, the encoder parameters are corrupted by the meaningless gradient through the untrained Linear, which makes the training unstable. To prevent this, I trained only the classifier first.</p>",
          "rawMarkdown": "Typically, we use imagenet pre-trained models for the encoder and untrained Linear for the classifier.　In the early stage of training, the encoder parameters are corrupted by the meaningless gradient through the untrained Linear, which makes the training unstable. To prevent this, I trained only the classifier first.",
          "votes": 7
        }
      ]
    },
    {
      "id": 1483145,
      "postDate": "2021-08-20T13:21:06.960Z",
      "content": "<p>Congratulations great job! I think your mixup, pseudo, and strong regularization was key!</p>\n<p>I've never used <code>StochasticDepth</code> before. Is it deployed with timm parameter <code>drop_path_rate</code>? Can it be applied to all timm models?</p>\n<pre><code>timm.create_model( drop_path_rate=RATE)\n</code></pre>",
      "rawMarkdown": "Congratulations great job! I think your mixup, pseudo, and strong regularization was key!\n\nI've never used `StochasticDepth` before. Is it deployed with timm parameter `drop_path_rate`? Can it be applied to all timm models?\n\n    timm.create_model( drop_path_rate=RATE)",
      "votes": 2,
      "replies": [
        {
          "id": 1483382,
          "postDate": "2021-08-20T15:45:31.533Z",
          "content": "<p>Yes, that's what I'm using. Apparently, it does not work for all models, but it does work for many models that are structured like ResNet.</p>",
          "rawMarkdown": "Yes, that's what I'm using. Apparently, it does not work for all models, but it does work for many models that are structured like ResNet.",
          "votes": 1
        }
      ]
    },
    {
      "id": 1486638,
      "postDate": "2021-08-23T05:48:55.310Z",
      "content": "<p>Congratulations! It seems like a simple probability operation can help improve prediction scores significantly when properly applied.  </p>",
      "rawMarkdown": "Congratulations! It seems like a simple probability operation can help improve prediction scores significantly when properly applied.  "
    },
    {
      "id": 1481719,
      "postDate": "2021-08-19T16:36:32.273Z",
      "content": "<p>Congrats and great summary! Thanks for sharing <a href=\"https://www.kaggle.com/hirune924\" target=\"_blank\">@hirune924</a> </p>",
      "rawMarkdown": "Congrats and great summary! Thanks for sharing @hirune924 "
    },
    {
      "id": 1480970,
      "postDate": "2021-08-19T08:51:30.243Z",
      "content": "<p>Congratulations on coming in second! And I want to a question about pseudo label.</p>\n<p>Is the data used in your second stage are training data and pseudo label? or only pseudo label?</p>",
      "rawMarkdown": "Congratulations on coming in second! And I want to a question about pseudo label.\n\nIs the data used in your second stage are training data and pseudo label? or only pseudo label?",
      "replies": [
        {
          "id": 1480997,
          "postDate": "2021-08-19T09:05:48.603Z",
          "content": "<p>Use both train and test.</p>",
          "rawMarkdown": "Use both train and test."
        },
        {
          "id": 1481090,
          "postDate": "2021-08-19T09:56:21.820Z",
          "content": "<p>Thanks for your answer. We do the same thing as your stage2 but we don't get much improve. It seems like regularizations are really important.</p>",
          "rawMarkdown": "Thanks for your answer. We do the same thing as your stage2 but we don't get much improve. It seems like regularizations are really important."
        }
      ]
    },
    {
      "id": 1480923,
      "postDate": "2021-08-19T08:20:11.807Z",
      "content": "<p>Congrats and thanks for sharing your solution!</p>",
      "rawMarkdown": "Congrats and thanks for sharing your solution!"
    },
    {
      "id": 1480761,
      "postDate": "2021-08-19T07:00:54.287Z",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/hirune924\" target=\"_blank\">@hirune924</a> can you tell us the model names that you have used in the final score?</p>",
      "rawMarkdown": "Hi @hirune924 can you tell us the model names that you have used in the final score?",
      "replies": [
        {
          "id": 1480785,
          "postDate": "2021-08-19T07:13:02.237Z",
          "content": "<p>Finally, the following models were blended. v2 was introduced to increase diversity.</p>\n<table>\n<thead>\n<tr>\n<th>model</th>\n<th>public LB</th>\n<th>private LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>tf_efficientnet_b5_ns</td>\n<td>0.81341</td>\n<td>0.81094</td>\n</tr>\n<tr>\n<td>tf_efficientnet_b5_ns</td>\n<td>0.81384</td>\n<td>0.81196</td>\n</tr>\n<tr>\n<td>tf_efficientnet_b5_ns</td>\n<td>0.81203</td>\n<td>0.81209</td>\n</tr>\n<tr>\n<td>tf_efficientnetv2_m_in21ft1k</td>\n<td>0.81169</td>\n<td>0.80978</td>\n</tr>\n<tr>\n<td>tf_efficientnetv2_m_in21ft1k</td>\n<td>0.80626</td>\n<td>0.80486</td>\n</tr>\n</tbody>\n</table>",
          "rawMarkdown": "Finally, the following models were blended. v2 was introduced to increase diversity.\n| model | public LB | private LB |\n| --- | --- | --- |\n| tf_efficientnet_b5_ns | 0.81341 | 0.81094 |\n| tf_efficientnet_b5_ns | 0.81384 | 0.81196 |\n| tf_efficientnet_b5_ns | 0.81203 | 0.81209 |\n| tf_efficientnetv2_m_in21ft1k| 0.81169 | 0.80978 |\n| tf_efficientnetv2_m_in21ft1k| 0.80626 | 0.80486 |",
          "votes": 5
        },
        {
          "id": 1481149,
          "postDate": "2021-08-19T10:43:26.887Z",
          "content": "<p>Thanks for sharing 😌</p>",
          "rawMarkdown": "Thanks for sharing 😌"
        }
      ]
    },
    {
      "id": 1480462,
      "postDate": "2021-08-19T03:21:11.483Z",
      "content": "<p>Congratulations!  That is a surprisingly simple solution. Thank you for sharing.</p>\n<p>I also used \"logical OR\" mixup, which is effective for CV and LB.<br>\nI didn't try pseudo labeling, I think that is a key.</p>\n<p>BTW, how did you use the \"logical OR\" mixup in training by pseudo label? by hard pseudo labeling?</p>",
      "rawMarkdown": "Congratulations!  That is a surprisingly simple solution. Thank you for sharing.\n\nI also used \"logical OR\" mixup, which is effective for CV and LB.\nI didn't try pseudo labeling, I think that is a key.\n\nBTW, how did you use the \"logical OR\" mixup in training by pseudo label? by hard pseudo labeling?",
      "replies": [
        {
          "id": 1480500,
          "postDate": "2021-08-19T03:52:07.677Z",
          "content": "<p>My method differs from strict logical OR, and works on soft targets because I used the following continuous method that works similarly.<br>\ny = y1 + y2 - (y1 * y2)</p>",
          "rawMarkdown": "My method differs from strict logical OR, and works on soft targets because I used the following continuous method that works similarly.\ny = y1 + y2 - (y1 * y2)",
          "votes": 1
        },
        {
          "id": 1480507,
          "postDate": "2021-08-19T03:56:58.150Z",
          "content": "<p>Thanks! I'll try it in late submission.</p>",
          "rawMarkdown": "Thanks! I'll try it in late submission."
        }
      ]
    },
    {
      "id": 1480456,
      "postDate": "2021-08-19T03:18:22.497Z",
      "content": "<p>Congratulations &amp; thanks for sharing the solution.</p>\n<p>I also tried logical-OR for the mixup target, but it didn't improve my CV score I wonder why.</p>\n<pre><code>y = torch.clamp(y_a + y_b, min=0, max=1)\n</code></pre>\n<p>Can you share us the 1st stage only score?<br>\nIn my experiments, I can't make LB score higher than 0.765, so I'm curious if the pseudo label is the key in the solution.</p>",
      "rawMarkdown": "Congratulations & thanks for sharing the solution.\n\nI also tried logical-OR for the mixup target, but it didn't improve my CV score I wonder why.\n```\ny = torch.clamp(y_a + y_b, min=0, max=1)\n```\n\nCan you share us the 1st stage only score?\nIn my experiments, I can't make LB score higher than 0.765, so I'm curious if the pseudo label is the key in the solution.",
      "replies": [
        {
          "id": 1480505,
          "postDate": "2021-08-19T03:56:04.863Z",
          "content": "<p>I achieve Public LB 0.80 at the first level alone. Also, the larger the model, the higher the score, but I reached the limit with b5!</p>",
          "rawMarkdown": "I achieve Public LB 0.80 at the first level alone. Also, the larger the model, the higher the score, but I reached the limit with b5!",
          "votes": 2
        },
        {
          "id": 1480554,
          "postDate": "2021-08-19T04:42:40.680Z",
          "content": "<p>So the model size are also important. I was stopping efficientnet-b2 due to computational resource limit. Thank you.</p>",
          "rawMarkdown": "So the model size are also important. I was stopping efficientnet-b2 due to computational resource limit. Thank you."
        }
      ]
    },
    {
      "id": 1480411,
      "postDate": "2021-08-19T02:43:24.133Z",
      "content": "<p>Thanks for sharing your approach!</p>\n<p>I am curious - how did you trained the b5 efficientnet on the 512x512 image size without overloading the GPU? Thanks for any help :)</p>",
      "rawMarkdown": "Thanks for sharing your approach!\n\nI am curious - how did you trained the b5 efficientnet on the 512x512 image size without overloading the GPU? Thanks for any help :)",
      "replies": [
        {
          "id": 1480425,
          "postDate": "2021-08-19T02:54:25.077Z",
          "content": "<p>V100x1 GPU is used to train b5, and mixed precision training is used to save GPU RAM.</p>",
          "rawMarkdown": "V100x1 GPU is used to train b5, and mixed precision training is used to save GPU RAM.",
          "votes": 2
        }
      ]
    },
    {
      "id": 1480366,
      "postDate": "2021-08-19T01:53:15.340Z",
      "content": "<p><a href=\"https://www.kaggle.com/hirune924\" target=\"_blank\">@hirune924</a> Congrats!🎉</p>\n<p>Thanks for sharing your approach. Interesting!</p>",
      "rawMarkdown": "@hirune924 Congrats!🎉\n\nThanks for sharing your approach. Interesting!"
    }
  ],
  "comments": [
    {
      "id": 1480685,
      "author_name": "Psi",
      "author_url": "",
      "post_date": "2021-08-19T06:17:21.457000",
      "content": "<p>Congrats, very impressive results. When we saw you climbing one point every day, we were quite convinced that you are pseudo tagging :)</p>\n<p>The important part here seems to do heavy regularization when incorporating the pseudo tags, so that the model does not overfit too heavily on the pseudo tags coming from images of a different distribution, we shortly discuss this also in our solution post. I believe that the best way to incorporate pseudo tags, in theory, would be to only add target=1 labels, because, at least from CV, the models make less mistakes when they predict and find the signal, compared to when they don't. However, this does not work here, because then the model would just learn that target=1 comes always from test (due to distribution shift). There might be some clever way to do this, by studying more the research on accounting for distribution shift (e.g., additional losses that try to prevent overfitting on the distribution). Unfortunately, didnt have time to look into it due to our other findings.</p>\n<p>Again: very impressive job and congrats on solo gold!</p>",
      "votes": 11,
      "replies": [
        {
          "id": 1480741,
          "author_name": "hirune924",
          "author_url": "",
          "post_date": "2021-08-19T06:48:19.727000",
          "content": "<p>Your team's solution is also a great job. Congratulations on the 1st place.</p>\n<p>Actually, I thought that using all the pseudo-labels instead of only target=1 is dangerous because it may overfit to the pseudo-labels.<br>\nTherefore, we tried to use only those pseudo-labels that showed extreme target scores by threshold, but the Noisy Student style method using all pseudo-labels showed better and more stable results.<br>\nI would have liked to increase the model size according to the paper, but there was not enough time.</p>",
          "votes": 5,
          "replies": []
        }
      ]
    },
    {
      "id": 1480428,
      "author_name": "Qishen Ha",
      "author_url": "",
      "post_date": "2021-08-19T02:56:21.657000",
      "content": "<p>Congratulations! your boost on 2nd stage is very inpressive!</p>\n<blockquote>\n  <p>y = y + y[index] - (y * y[index])</p>\n</blockquote>\n<p>This is completely the same as our mixup: <code>torch.stack([y, y[index]], 0).max(0).values</code></p>\n<p>But we didn't adopted PL or NS…</p>",
      "votes": 5,
      "replies": [
        {
          "id": 1480737,
          "author_name": "Mr_KnowNothing",
          "author_url": "",
          "post_date": "2021-08-19T06:45:27.797000",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/haqishen\" target=\"_blank\">@haqishen</a> my intuition was that if we use mixup it will only distort the signal and will not be good . I am still very weak in CV can you please tell us the thought process(intuition) behind using mixup and the trick used above</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1481066,
          "author_name": "Qishen Ha",
          "author_url": "",
          "post_date": "2021-08-19T09:42:49.290000",
          "content": "<p>Actually weaken the injected signal is what we want, because we need the model to gain the ability to detect weak signal as well. The modified mixup is base on this intuition.</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 1481145,
          "author_name": "Mr_KnowNothing",
          "author_url": "",
          "post_date": "2021-08-19T10:42:21.307000",
          "content": "<p>Thank you </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1480397,
      "author_name": "CPMP",
      "author_url": "",
      "post_date": "2021-08-19T02:26:27.783000",
      "content": "<p>Congrats!  Your solution looks quite similar to ours, but you gain more from pseudo labeling test.  I wonder why. Interestingly I used a very similar trick for target in mixup:</p>\n<pre><code>            target = alpha * target + (1 - alpha) * target[perm]\n            target = torch.clamp(2*target, 0, 1)\n</code></pre>",
      "votes": 5,
      "replies": [
        {
          "id": 1480430,
          "author_name": "hirune924",
          "author_url": "",
          "post_date": "2021-08-19T02:56:38.207000",
          "content": "<p>I'm honored to have arrived at the same trick as you !</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1481175,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2021-08-19T11:02:06.487000",
          "content": "<p>I am the one who is honored given you did much better than us!</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1482292,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-08-20T02:23:21.267000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1482342,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2021-08-20T03:26:46.580000",
          "content": "<p>let say alpha is 0.5 i.e. the new image is the mean of the two mixed images. If one of them is of target 1 and the other 0, the mean target is 0.5.  By multiplying by 2 it becomes 1=, which is what we want here.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1482385,
          "author_name": "Hao",
          "author_url": "",
          "post_date": "2021-08-20T04:00:37.020000",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> , regarding how to mix the target value during Mixup, sometimes people use standard(weighted) way like original <a href=\"https://arxiv.org/pdf/1710.09412v2.pdf\" target=\"_blank\">paper</a>, sometimes use maximum(logical OR) way like you did, I'm wondering what is the benefit for each way? The latter way is better for datasets which have very few positive samples, or target singal is very weak like this competition?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1483260,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2021-08-20T14:36:55.963000",
          "content": "<p>Here we look for messages.  A message stays in the mixed image, and the hope is that this way model will learn to recognized weaker signals (weaker because they are blended with another image).</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1480373,
      "author_name": "patriot",
      "author_url": "",
      "post_date": "2021-08-19T02:04:17.630000",
      "content": "<p>Thanks for sharing!!　Congratulations on the solo gold medal!!<br>\nI also tried arcface(m=0.5) to cope with unseen cls, but it failed. it takes 2x times than normal training and score wasn't good(LB=0.776,B0, size=768<em>768</em>3 )</p>",
      "votes": 3,
      "replies": [
        {
          "id": 1480434,
          "author_name": "hirune924",
          "author_url": "",
          "post_date": "2021-08-19T02:59:06.740000",
          "content": "<p>I tried L2 constrained softmax Loss, but the max I could get was about lb0.73.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1495846,
      "author_name": "Fuco",
      "author_url": "",
      "post_date": "2021-08-29T20:40:49.727000",
      "content": "<p>Great work! Congrats</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1480504,
      "author_name": "Hao",
      "author_url": "",
      "post_date": "2021-08-19T03:56:00.087000",
      "content": "<p>Congratulations for the solo gold <a href=\"https://www.kaggle.com/hirune924\" target=\"_blank\">@hirune924</a>  🎉</p>\n<blockquote>\n  <p>the loss became nan or the validation auc became 0.5 and the training did not progress. In this case, I solved the problem by training only the head and encoder BNs for one epoch first.</p>\n</blockquote>\n<p>I also observed sometimes my model does not learn anything as validation AUC is alway ~0.5, especially when I add more augmentations like ShiftScaleRotate or RandomResizedCrop, finally I found switch to OneCycleLR scheduler would help a bit, I suspect it`s due to it will start training with a small LR(1e-6 eg.) to find corrent direction to converge. May I know what is your theory about this \"always 0.5 CV\" problem? And why training with only head and encoder BNs in 1st epoch would help? Thanks.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1480528,
          "author_name": "hirune924",
          "author_url": "",
          "post_date": "2021-08-19T04:16:26.527000",
          "content": "<p>Typically, we use imagenet pre-trained models for the encoder and untrained Linear for the classifier.　In the early stage of training, the encoder parameters are corrupted by the meaningless gradient through the untrained Linear, which makes the training unstable. To prevent this, I trained only the classifier first.</p>",
          "votes": 7,
          "replies": []
        }
      ]
    },
    {
      "id": 1483145,
      "author_name": "Chris Deotte",
      "author_url": "",
      "post_date": "2021-08-20T13:21:06.960000",
      "content": "<p>Congratulations great job! I think your mixup, pseudo, and strong regularization was key!</p>\n<p>I've never used <code>StochasticDepth</code> before. Is it deployed with timm parameter <code>drop_path_rate</code>? Can it be applied to all timm models?</p>\n<pre><code>timm.create_model( drop_path_rate=RATE)\n</code></pre>",
      "votes": 2,
      "replies": [
        {
          "id": 1483382,
          "author_name": "hirune924",
          "author_url": "",
          "post_date": "2021-08-20T15:45:31.533000",
          "content": "<p>Yes, that's what I'm using. Apparently, it does not work for all models, but it does work for many models that are structured like ResNet.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1486638,
      "author_name": "Dac-Thanh Van",
      "author_url": "",
      "post_date": "2021-08-23T05:48:55.310000",
      "content": "<p>Congratulations! It seems like a simple probability operation can help improve prediction scores significantly when properly applied.  </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1481719,
      "author_name": "Kalilur Rahman",
      "author_url": "",
      "post_date": "2021-08-19T16:36:32.273000",
      "content": "<p>Congrats and great summary! Thanks for sharing <a href=\"https://www.kaggle.com/hirune924\" target=\"_blank\">@hirune924</a> </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1480970,
      "author_name": "Matthew Wu",
      "author_url": "",
      "post_date": "2021-08-19T08:51:30.243000",
      "content": "<p>Congratulations on coming in second! And I want to a question about pseudo label.</p>\n<p>Is the data used in your second stage are training data and pseudo label? or only pseudo label?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1480997,
          "author_name": "hirune924",
          "author_url": "",
          "post_date": "2021-08-19T09:05:48.603000",
          "content": "<p>Use both train and test.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1481090,
          "author_name": "Matthew Wu",
          "author_url": "",
          "post_date": "2021-08-19T09:56:21.820000",
          "content": "<p>Thanks for your answer. We do the same thing as your stage2 but we don't get much improve. It seems like regularizations are really important.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1480923,
      "author_name": "Cuong Thanh-Viet Nguyen",
      "author_url": "",
      "post_date": "2021-08-19T08:20:11.807000",
      "content": "<p>Congrats and thanks for sharing your solution!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1480761,
      "author_name": "Mr_KnowNothing",
      "author_url": "",
      "post_date": "2021-08-19T07:00:54.287000",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/hirune924\" target=\"_blank\">@hirune924</a> can you tell us the model names that you have used in the final score?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1480785,
          "author_name": "hirune924",
          "author_url": "",
          "post_date": "2021-08-19T07:13:02.237000",
          "content": "<p>Finally, the following models were blended. v2 was introduced to increase diversity.</p>\n<table>\n<thead>\n<tr>\n<th>model</th>\n<th>public LB</th>\n<th>private LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>tf_efficientnet_b5_ns</td>\n<td>0.81341</td>\n<td>0.81094</td>\n</tr>\n<tr>\n<td>tf_efficientnet_b5_ns</td>\n<td>0.81384</td>\n<td>0.81196</td>\n</tr>\n<tr>\n<td>tf_efficientnet_b5_ns</td>\n<td>0.81203</td>\n<td>0.81209</td>\n</tr>\n<tr>\n<td>tf_efficientnetv2_m_in21ft1k</td>\n<td>0.81169</td>\n<td>0.80978</td>\n</tr>\n<tr>\n<td>tf_efficientnetv2_m_in21ft1k</td>\n<td>0.80626</td>\n<td>0.80486</td>\n</tr>\n</tbody>\n</table>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 1481149,
          "author_name": "Mr_KnowNothing",
          "author_url": "",
          "post_date": "2021-08-19T10:43:26.887000",
          "content": "<p>Thanks for sharing 😌</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1480462,
      "author_name": "Tawara",
      "author_url": "",
      "post_date": "2021-08-19T03:21:11.483000",
      "content": "<p>Congratulations!  That is a surprisingly simple solution. Thank you for sharing.</p>\n<p>I also used \"logical OR\" mixup, which is effective for CV and LB.<br>\nI didn't try pseudo labeling, I think that is a key.</p>\n<p>BTW, how did you use the \"logical OR\" mixup in training by pseudo label? by hard pseudo labeling?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1480500,
          "author_name": "hirune924",
          "author_url": "",
          "post_date": "2021-08-19T03:52:07.677000",
          "content": "<p>My method differs from strict logical OR, and works on soft targets because I used the following continuous method that works similarly.<br>\ny = y1 + y2 - (y1 * y2)</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1480507,
          "author_name": "Tawara",
          "author_url": "",
          "post_date": "2021-08-19T03:56:58.150000",
          "content": "<p>Thanks! I'll try it in late submission.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1480456,
      "author_name": "Bilzard",
      "author_url": "",
      "post_date": "2021-08-19T03:18:22.497000",
      "content": "<p>Congratulations &amp; thanks for sharing the solution.</p>\n<p>I also tried logical-OR for the mixup target, but it didn't improve my CV score I wonder why.</p>\n<pre><code>y = torch.clamp(y_a + y_b, min=0, max=1)\n</code></pre>\n<p>Can you share us the 1st stage only score?<br>\nIn my experiments, I can't make LB score higher than 0.765, so I'm curious if the pseudo label is the key in the solution.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1480505,
          "author_name": "hirune924",
          "author_url": "",
          "post_date": "2021-08-19T03:56:04.863000",
          "content": "<p>I achieve Public LB 0.80 at the first level alone. Also, the larger the model, the higher the score, but I reached the limit with b5!</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1480554,
          "author_name": "Bilzard",
          "author_url": "",
          "post_date": "2021-08-19T04:42:40.680000",
          "content": "<p>So the model size are also important. I was stopping efficientnet-b2 due to computational resource limit. Thank you.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1480411,
      "author_name": "Daniel Furman",
      "author_url": "",
      "post_date": "2021-08-19T02:43:24.133000",
      "content": "<p>Thanks for sharing your approach!</p>\n<p>I am curious - how did you trained the b5 efficientnet on the 512x512 image size without overloading the GPU? Thanks for any help :)</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1480425,
          "author_name": "hirune924",
          "author_url": "",
          "post_date": "2021-08-19T02:54:25.077000",
          "content": "<p>V100x1 GPU is used to train b5, and mixed precision training is used to save GPU RAM.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 1480366,
      "author_name": "BaAis",
      "author_url": "",
      "post_date": "2021-08-19T01:53:15.340000",
      "content": "<p><a href=\"https://www.kaggle.com/hirune924\" target=\"_blank\">@hirune924</a> Congrats!🎉</p>\n<p>Thanks for sharing your approach. Interesting!</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1480362": "Thanks to Kaggle and Berkeley SETI Research Center for this competition. Finding the aliens was very exciting. I tried many things to close the gap between cv and lb, but few succeeded, and as a result, my solution is very simple.\n\n## Abstract\nMy solution is a two-step process. The first stage is training with train data only, and the next stage is training with test data and pseudo labels. In the first stage, I use only vflip, cutout, and mixup to prevent the model from being confused by losing the signal due to unexpected augmentation. In addition, StochasticDepth and Dropout are strongly used to prevent overtraining. The mixup uses logical OR instead of alpha blending when mixing targets. This allows creating a model that responds strongly to weak signals. At this stage, we achieved PublicLB 0.800 by training efficientnetB5.\n\nIn the second stage, we finetune the first stage model using pseudo-labels created by using the pre-trained model from the first stage. In this case, we refer to NoisyStudent and use Stochastic Depth and Dropout more strongly to prevent overfitting to the pseudo-label. By repeating this process several times, the PublicLB 0.813 is reached. Finally, I simply averaged the five submissions with the highest LB scores.\n\n## 1st Stage Details\nYou can see a sample of the code here.\nhttps://www.kaggle.com/hirune924/2ndplace-solution\n* The input images are merged in the time direction and resized to a size of 512x512.\n* The mixup target can be mixed by using the following expression to express a logical OR, which also supports soft targets when using pseudo labels.\n```\ny = y + y[index] - (y * y[index])\n```\n* TTA with vflip is used for inference\n\nSometimes, the loss became nan or the validation auc became 0.5 and the training did not progress. In this case, I solved the problem by training only the head and encoder BNs for one epoch first.\n\n## 2nd Stage Details\nThe code is almost the same as the first step, adding pseudo-labels and making Stochastic Depth and Dropout even stronger for training.\n\n## Didn't work\n* Unsupervised Domain Adaptation by finetune only BN\n* Add artificial signal, chi2noise, referring to setigen.\n* input stem stride=1\n* using old data\n* Anomaly detection based metric learning\n* Semi Supervised Learning\n* many other things...\n\n",
    "1480685": "Congrats, very impressive results. When we saw you climbing one point every day, we were quite convinced that you are pseudo tagging :)\n\nThe important part here seems to do heavy regularization when incorporating the pseudo tags, so that the model does not overfit too heavily on the pseudo tags coming from images of a different distribution, we shortly discuss this also in our solution post. I believe that the best way to incorporate pseudo tags, in theory, would be to only add target=1 labels, because, at least from CV, the models make less mistakes when they predict and find the signal, compared to when they don't. However, this does not work here, because then the model would just learn that target=1 comes always from test (due to distribution shift). There might be some clever way to do this, by studying more the research on accounting for distribution shift (e.g., additional losses that try to prevent overfitting on the distribution). Unfortunately, didnt have time to look into it due to our other findings.\n\nAgain: very impressive job and congrats on solo gold!",
    "1480428": "Congratulations! your boost on 2nd stage is very inpressive!\n\n> y = y + y[index] - (y * y[index])\n\nThis is completely the same as our mixup: `torch.stack([y, y[index]], 0).max(0).values`\n\nBut we didn't adopted PL or NS...",
    "1480397": "Congrats!  Your solution looks quite similar to ours, but you gain more from pseudo labeling test.  I wonder why. Interestingly I used a very similar trick for target in mixup:\n\n                target = alpha * target + (1 - alpha) * target[perm]\n                target = torch.clamp(2*target, 0, 1)\n",
    "1480373": "Thanks for sharing!!　Congratulations on the solo gold medal!!\nI also tried arcface(m=0.5) to cope with unseen cls, but it failed. it takes 2x times than normal training and score wasn't good(LB=0.776,B0, size=768*768*3 )",
    "1495846": "Great work! Congrats",
    "1480504": "Congratulations for the solo gold @hirune924  🎉\n\n> the loss became nan or the validation auc became 0.5 and the training did not progress. In this case, I solved the problem by training only the head and encoder BNs for one epoch first.\n\nI also observed sometimes my model does not learn anything as validation AUC is alway ~0.5, especially when I add more augmentations like ShiftScaleRotate or RandomResizedCrop, finally I found switch to OneCycleLR scheduler would help a bit, I suspect it`s due to it will start training with a small LR(1e-6 eg.) to find corrent direction to converge. May I know what is your theory about this \"always 0.5 CV\" problem? And why training with only head and encoder BNs in 1st epoch would help? Thanks.",
    "1483145": "Congratulations great job! I think your mixup, pseudo, and strong regularization was key!\n\nI've never used `StochasticDepth` before. Is it deployed with timm parameter `drop_path_rate`? Can it be applied to all timm models?\n\n    timm.create_model( drop_path_rate=RATE)",
    "1486638": "Congratulations! It seems like a simple probability operation can help improve prediction scores significantly when properly applied.  ",
    "1481719": "Congrats and great summary! Thanks for sharing @hirune924 ",
    "1480970": "Congratulations on coming in second! And I want to a question about pseudo label.\n\nIs the data used in your second stage are training data and pseudo label? or only pseudo label?",
    "1480923": "Congrats and thanks for sharing your solution!",
    "1480761": "Hi @hirune924 can you tell us the model names that you have used in the final score?",
    "1480462": "Congratulations!  That is a surprisingly simple solution. Thank you for sharing.\n\nI also used \"logical OR\" mixup, which is effective for CV and LB.\nI didn't try pseudo labeling, I think that is a key.\n\nBTW, how did you use the \"logical OR\" mixup in training by pseudo label? by hard pseudo labeling?",
    "1480456": "Congratulations & thanks for sharing the solution.\n\nI also tried logical-OR for the mixup target, but it didn't improve my CV score I wonder why.\n```\ny = torch.clamp(y_a + y_b, min=0, max=1)\n```\n\nCan you share us the 1st stage only score?\nIn my experiments, I can't make LB score higher than 0.765, so I'm curious if the pseudo label is the key in the solution.",
    "1480411": "Thanks for sharing your approach!\n\nI am curious - how did you trained the b5 efficientnet on the 512x512 image size without overloading the GPU? Thanks for any help :)",
    "1480366": "@hirune924 Congrats!🎉\n\nThanks for sharing your approach. Interesting!"
  }
}