{
  "id": 266403,
  "title": "3rd place solution (Triplet Attention)",
  "url": "/competitions/seti-breakthrough-listen/discussion/266403",
  "author_name": "kenji",
  "post_date": "2021-08-19T02:57:35.455000",
  "votes": 65,
  "comment_count": 26,
  "views": 0,
  "content": "<p>We thank all organizers for this very exciting competition.  <br>\nCongratulations to all who finished the competition and to the winners.</p>\n<h2>Summary</h2>\n<p>Best submission was an emsemble of four EfficientNets with <a href=\"https://arxiv.org/abs/2010.03045\" target=\"_blank\">Convolutional Triplet Attention Module</a>.</p>\n<p><strong>This post was updated on September 1, 2021.</strong><br>\nThe main changes are as follows.</p>\n<ol>\n<li>Added stage0 to train the model for pseudo-label creation.</li>\n<li>Already used pseudo-labels in stage1, and fixed that.</li>\n<li>Update initial learning rate value.</li>\n</ol>\n<p>The implementation, including training code, has also been released at the following URL.<br>\n<a href=\"https://github.com/knjcode/kaggle-seti-2021\" target=\"_blank\">https://github.com/knjcode/kaggle-seti-2021</a></p>\n<h2>Validation and Preprocess</h2>\n<ul>\n<li>Use new train data only</li>\n<li>StratifiedKFold(k=5)</li>\n<li>Use ON-channels only (819x256) and resize (768x768)</li>\n</ul>\n<h2>Model Architecture</h2>\n<p>Backbone -&gt; Triplet Attention -&gt; GeM Pooling -&gt; FC</p>\n<ul>\n<li>Use <a href=\"https://arxiv.org/abs/1905.11946\" target=\"_blank\">EfficientNet</a> B1, B2, B3 and B4 backbone</li>\n<li><a href=\"https://arxiv.org/abs/2010.03045\" target=\"_blank\">Triplet Attention</a> (increase kernel size from 7 to 13)<ul>\n<li>Add only one attention layer after the backbone</li>\n<li>I used the implementation in <a href=\"https://github.com/rwightman/pytorch-image-models/blob/499790e117b2c8c1b57780b73d16c28b84db509e/timm/models/layers/triplet.py\" target=\"_blank\">here</a></li></ul></li>\n<li><a href=\"https://arxiv.org/abs/1711.02512\" target=\"_blank\">GeM Pooling</a> (p=4)</li>\n<li>To accelerate the training, replace Swish to ReLU (B3 and B4 only)</li>\n</ul>\n<h2>Training</h2>\n<p>The training process consists of three stages</p>\n<h3>Stage 0</h3>\n<p>This model will not be used for the final submission.</p>\n<p>Training EfficientNet-B4</p>\n<ul>\n<li>60epochs DDP AMP</li>\n<li>Focal Loss (gamma=0.5)</li>\n<li><a href=\"https://github.com/facebookresearch/madgrad\" target=\"_blank\">MADGRAD</a> optimizer<ul>\n<li>Initial Lr 1e-2 LinearWarmupCosineAnnealingLR (warmup_epochs=1)</li></ul></li>\n<li>Mixup (alpha=1.0)<ul>\n<li>If the target of either data is 1, mixed_target will also be set to 1</li></ul></li>\n<li>Data Augmentation<ul>\n<li>Horizontal and Vertical flip (p=0.5)</li>\n<li>ShiftScaleRotate (shift_limit=0.2, scale_limit=0.2, rotate_limit=0)</li>\n<li>RandomResizedCrop (scale=[0.8,1.0], ratio=[0.75,1.333])</li></ul></li>\n</ul>\n<p>Generate Pseudo labels from oof prediction of this model from new test set.  <br>\n(Select 5,000 images each of positive and negative data with high confidence)</p>\n<h3>Stage 1</h3>\n<p>Training EfficientNet-B1,B2,B3 and B4 by adding pseudo labeled images.</p>\n<p>The training settings are almost the same as stage0.</p>\n<ul>\n<li>60epochs DDP AMP</li>\n<li>Focal Loss (gamma=0.5)</li>\n<li><a href=\"https://github.com/facebookresearch/madgrad\" target=\"_blank\">MADGRAD</a> optimizer<ul>\n<li>Initial Lr 1e-2 or 1e-3 LinearWarmupCosineAnnealingLR (warmup_epochs=5)</li></ul></li>\n<li>Mixup (alpha=1.0)<ul>\n<li>If the target of either data is 1, mixed_target will also be set to 1</li></ul></li>\n<li>Data Augmentation<ul>\n<li>Horizontal and Vertical flip (p=0.5)</li>\n<li>ShiftScaleRotate (shift_limit=0.2, scale_limit=0.2, rotate_limit=0)</li>\n<li>RandomResizedCrop (scale=[0.8,1.0], ratio=[0.75,1.333])</li></ul></li>\n</ul>\n<h3>Stage2</h3>\n<p>Refine stage1 models and generate submission with TTA</p>\n<p>I refined the model with lighter data augmentation than stage1 for 10epochs.</p>\n<ul>\n<li>10epochs DDP AMP</li>\n<li>Focal Loss (gamma=0.5)</li>\n<li>MADGRAD optimizer<ul>\n<li>Initial Lr 1e-4 LinearWarmupCosineAnnealingLR (warmup_epochs=0)</li></ul></li>\n<li>Horizontal and Vertical flip only with Mixup</li>\n<li>4xTTA in inference<ul>\n<li>regular, hflip, vflip, hflip+vflip</li></ul></li>\n</ul>\n<p>Best Single Model  (Efficientnet-B4)</p>\n<ul>\n<li>5foldCV: 0.9048   PublicLB: 0.80586   PrivateLB: 0.80294</li>\n</ul>\n<p>Ensemble of 4 EfficinetNets</p>\n<ul>\n<li>5foldCV: 0.9100   PublicLB: 0.80575   PrivateLB: 0.80475</li>\n</ul>\n<h2>Reference</h2>\n<ul>\n<li><a href=\"https://arxiv.org/abs/1905.11946\" target=\"_blank\">EfficientNet: Rethinking Model Scaling for Convolutional Neural Networks</a></li>\n<li><a href=\"https://arxiv.org/abs/2010.03045\" target=\"_blank\">Rotate to Attend: Convolutional Triplet Attention Module</a></li>\n<li><a href=\"https://arxiv.org/abs/1711.02512\" target=\"_blank\">Fine-tuning CNN Image Retrieval with No Human Annotation</a></li>\n</ul>",
  "messages": [
    {
      "id": 1480432,
      "postDate": "2021-08-19T02:57:35.457Z",
      "content": "<p>We thank all organizers for this very exciting competition.  <br>\nCongratulations to all who finished the competition and to the winners.</p>\n<h2>Summary</h2>\n<p>Best submission was an emsemble of four EfficientNets with <a href=\"https://arxiv.org/abs/2010.03045\" target=\"_blank\">Convolutional Triplet Attention Module</a>.</p>\n<p><strong>This post was updated on September 1, 2021.</strong><br>\nThe main changes are as follows.</p>\n<ol>\n<li>Added stage0 to train the model for pseudo-label creation.</li>\n<li>Already used pseudo-labels in stage1, and fixed that.</li>\n<li>Update initial learning rate value.</li>\n</ol>\n<p>The implementation, including training code, has also been released at the following URL.<br>\n<a href=\"https://github.com/knjcode/kaggle-seti-2021\" target=\"_blank\">https://github.com/knjcode/kaggle-seti-2021</a></p>\n<h2>Validation and Preprocess</h2>\n<ul>\n<li>Use new train data only</li>\n<li>StratifiedKFold(k=5)</li>\n<li>Use ON-channels only (819x256) and resize (768x768)</li>\n</ul>\n<h2>Model Architecture</h2>\n<p>Backbone -&gt; Triplet Attention -&gt; GeM Pooling -&gt; FC</p>\n<ul>\n<li>Use <a href=\"https://arxiv.org/abs/1905.11946\" target=\"_blank\">EfficientNet</a> B1, B2, B3 and B4 backbone</li>\n<li><a href=\"https://arxiv.org/abs/2010.03045\" target=\"_blank\">Triplet Attention</a> (increase kernel size from 7 to 13)<ul>\n<li>Add only one attention layer after the backbone</li>\n<li>I used the implementation in <a href=\"https://github.com/rwightman/pytorch-image-models/blob/499790e117b2c8c1b57780b73d16c28b84db509e/timm/models/layers/triplet.py\" target=\"_blank\">here</a></li></ul></li>\n<li><a href=\"https://arxiv.org/abs/1711.02512\" target=\"_blank\">GeM Pooling</a> (p=4)</li>\n<li>To accelerate the training, replace Swish to ReLU (B3 and B4 only)</li>\n</ul>\n<h2>Training</h2>\n<p>The training process consists of three stages</p>\n<h3>Stage 0</h3>\n<p>This model will not be used for the final submission.</p>\n<p>Training EfficientNet-B4</p>\n<ul>\n<li>60epochs DDP AMP</li>\n<li>Focal Loss (gamma=0.5)</li>\n<li><a href=\"https://github.com/facebookresearch/madgrad\" target=\"_blank\">MADGRAD</a> optimizer<ul>\n<li>Initial Lr 1e-2 LinearWarmupCosineAnnealingLR (warmup_epochs=1)</li></ul></li>\n<li>Mixup (alpha=1.0)<ul>\n<li>If the target of either data is 1, mixed_target will also be set to 1</li></ul></li>\n<li>Data Augmentation<ul>\n<li>Horizontal and Vertical flip (p=0.5)</li>\n<li>ShiftScaleRotate (shift_limit=0.2, scale_limit=0.2, rotate_limit=0)</li>\n<li>RandomResizedCrop (scale=[0.8,1.0], ratio=[0.75,1.333])</li></ul></li>\n</ul>\n<p>Generate Pseudo labels from oof prediction of this model from new test set.  <br>\n(Select 5,000 images each of positive and negative data with high confidence)</p>\n<h3>Stage 1</h3>\n<p>Training EfficientNet-B1,B2,B3 and B4 by adding pseudo labeled images.</p>\n<p>The training settings are almost the same as stage0.</p>\n<ul>\n<li>60epochs DDP AMP</li>\n<li>Focal Loss (gamma=0.5)</li>\n<li><a href=\"https://github.com/facebookresearch/madgrad\" target=\"_blank\">MADGRAD</a> optimizer<ul>\n<li>Initial Lr 1e-2 or 1e-3 LinearWarmupCosineAnnealingLR (warmup_epochs=5)</li></ul></li>\n<li>Mixup (alpha=1.0)<ul>\n<li>If the target of either data is 1, mixed_target will also be set to 1</li></ul></li>\n<li>Data Augmentation<ul>\n<li>Horizontal and Vertical flip (p=0.5)</li>\n<li>ShiftScaleRotate (shift_limit=0.2, scale_limit=0.2, rotate_limit=0)</li>\n<li>RandomResizedCrop (scale=[0.8,1.0], ratio=[0.75,1.333])</li></ul></li>\n</ul>\n<h3>Stage2</h3>\n<p>Refine stage1 models and generate submission with TTA</p>\n<p>I refined the model with lighter data augmentation than stage1 for 10epochs.</p>\n<ul>\n<li>10epochs DDP AMP</li>\n<li>Focal Loss (gamma=0.5)</li>\n<li>MADGRAD optimizer<ul>\n<li>Initial Lr 1e-4 LinearWarmupCosineAnnealingLR (warmup_epochs=0)</li></ul></li>\n<li>Horizontal and Vertical flip only with Mixup</li>\n<li>4xTTA in inference<ul>\n<li>regular, hflip, vflip, hflip+vflip</li></ul></li>\n</ul>\n<p>Best Single Model  (Efficientnet-B4)</p>\n<ul>\n<li>5foldCV: 0.9048   PublicLB: 0.80586   PrivateLB: 0.80294</li>\n</ul>\n<p>Ensemble of 4 EfficinetNets</p>\n<ul>\n<li>5foldCV: 0.9100   PublicLB: 0.80575   PrivateLB: 0.80475</li>\n</ul>\n<h2>Reference</h2>\n<ul>\n<li><a href=\"https://arxiv.org/abs/1905.11946\" target=\"_blank\">EfficientNet: Rethinking Model Scaling for Convolutional Neural Networks</a></li>\n<li><a href=\"https://arxiv.org/abs/2010.03045\" target=\"_blank\">Rotate to Attend: Convolutional Triplet Attention Module</a></li>\n<li><a href=\"https://arxiv.org/abs/1711.02512\" target=\"_blank\">Fine-tuning CNN Image Retrieval with No Human Annotation</a></li>\n</ul>",
      "rawMarkdown": "We thank all organizers for this very exciting competition.  \nCongratulations to all who finished the competition and to the winners.\n\n## Summary\n\nBest submission was an emsemble of four EfficientNets with [Convolutional Triplet Attention Module](https://arxiv.org/abs/2010.03045).\n\n**This post was updated on September 1, 2021.**\nThe main changes are as follows.\n1. Added stage0 to train the model for pseudo-label creation.\n2. Already used pseudo-labels in stage1, and fixed that.\n3. Update initial learning rate value.\n\nThe implementation, including training code, has also been released at the following URL.\nhttps://github.com/knjcode/kaggle-seti-2021\n\n\n## Validation and Preprocess\n\n- Use new train data only\n- StratifiedKFold(k=5)\n- Use ON-channels only (819x256) and resize (768x768)\n\n\n## Model Architecture\n\nBackbone -> Triplet Attention -> GeM Pooling -> FC\n\n- Use [EfficientNet](https://arxiv.org/abs/1905.11946) B1, B2, B3 and B4 backbone\n- [Triplet Attention](https://arxiv.org/abs/2010.03045) (increase kernel size from 7 to 13)\n  - Add only one attention layer after the backbone\n  - I used the implementation in [here](https://github.com/rwightman/pytorch-image-models/blob/499790e117b2c8c1b57780b73d16c28b84db509e/timm/models/layers/triplet.py)\n- [GeM Pooling](https://arxiv.org/abs/1711.02512) (p=4)\n- To accelerate the training, replace Swish to ReLU (B3 and B4 only)\n\n\n## Training\n\nThe training process consists of three stages\n\n### Stage 0\n\nThis model will not be used for the final submission.\n\nTraining EfficientNet-B4\n\n- 60epochs DDP AMP\n- Focal Loss (gamma=0.5)\n- [MADGRAD](https://github.com/facebookresearch/madgrad) optimizer\n  - Initial Lr 1e-2 LinearWarmupCosineAnnealingLR (warmup_epochs=1)\n- Mixup (alpha=1.0)\n  - If the target of either data is 1, mixed_target will also be set to 1\n- Data Augmentation\n  - Horizontal and Vertical flip (p=0.5)\n  - ShiftScaleRotate (shift_limit=0.2, scale_limit=0.2, rotate_limit=0)\n  - RandomResizedCrop (scale=[0.8,1.0], ratio=[0.75,1.333])\n\nGenerate Pseudo labels from oof prediction of this model from new test set.  \n(Select 5,000 images each of positive and negative data with high confidence)\n\n### Stage 1\n\nTraining EfficientNet-B1,B2,B3 and B4 by adding pseudo labeled images.\n\nThe training settings are almost the same as stage0.\n\n- 60epochs DDP AMP\n- Focal Loss (gamma=0.5)\n- [MADGRAD](https://github.com/facebookresearch/madgrad) optimizer\n  - Initial Lr 1e-2 or 1e-3 LinearWarmupCosineAnnealingLR (warmup_epochs=5)\n- Mixup (alpha=1.0)\n  - If the target of either data is 1, mixed_target will also be set to 1\n- Data Augmentation\n  - Horizontal and Vertical flip (p=0.5)\n  - ShiftScaleRotate (shift_limit=0.2, scale_limit=0.2, rotate_limit=0)\n  - RandomResizedCrop (scale=[0.8,1.0], ratio=[0.75,1.333])\n\n\n### Stage2\n\nRefine stage1 models and generate submission with TTA\n\nI refined the model with lighter data augmentation than stage1 for 10epochs.\n\n- 10epochs DDP AMP\n- Focal Loss (gamma=0.5)\n- MADGRAD optimizer\n  - Initial Lr 1e-4 LinearWarmupCosineAnnealingLR (warmup_epochs=0)\n- Horizontal and Vertical flip only with Mixup\n- 4xTTA in inference\n  - regular, hflip, vflip, hflip+vflip\n\n\nBest Single Model  (Efficientnet-B4)\n- 5foldCV: 0.9048   PublicLB: 0.80586   PrivateLB: 0.80294\n\nEnsemble of 4 EfficinetNets\n- 5foldCV: 0.9100   PublicLB: 0.80575   PrivateLB: 0.80475\n\n\n## Reference\n\n- [EfficientNet: Rethinking Model Scaling for Convolutional Neural Networks](https://arxiv.org/abs/1905.11946)\n- [Rotate to Attend: Convolutional Triplet Attention Module](https://arxiv.org/abs/2010.03045)\n- [Fine-tuning CNN Image Retrieval with No Human Annotation](https://arxiv.org/abs/1711.02512)\n",
      "votes": 65
    },
    {
      "id": 1495844,
      "postDate": "2021-08-29T20:40:06.090Z",
      "content": "<p>Congratulations &amp; thanks for sharing the solution.</p>",
      "rawMarkdown": "Congratulations & thanks for sharing the solution.\n\n",
      "votes": 1
    },
    {
      "id": 1484109,
      "postDate": "2021-08-21T04:35:13.460Z",
      "content": "<p>Congratulations </p>",
      "rawMarkdown": "Congratulations "
    },
    {
      "id": 1483001,
      "postDate": "2021-08-20T11:47:09.177Z",
      "content": "<p>Congrats! It is very interesting.</p>",
      "rawMarkdown": "Congrats! It is very interesting."
    },
    {
      "id": 1482729,
      "postDate": "2021-08-20T08:14:56.077Z",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/knjcode\" target=\"_blank\">@knjcode</a> , May i ask how you used the extra class(s-shape target that only in test data) for training,did you added it to training set as extra data or used it as augmentation and set the target to 1 ?  if you manully added to training set, how many samples did you add? </p>",
      "rawMarkdown": "Congratulations @knjcode , May i ask how you used the extra class(s-shape target that only in test data) for training,did you added it to training set as extra data or used it as augmentation and set the target to 1 ?  if you manully added to training set, how many samples did you add? ",
      "replies": [
        {
          "id": 1482793,
          "postDate": "2021-08-20T09:29:09.163Z",
          "content": "<p>Thanks!</p>\n<p>I have trained the model using only the new train set.</p>\n<p>However, since I used pseudo labels to train the model in second stage training.<br>\nIf there is s-shape target that first stage model can determine with high confidence (if it is in top 5,000 positive), it was included in the training data.</p>\n<p>Also, I use soft target for pseudo labels.<br>\nIt is not the value of 0,1, but the value of the probability predicted by the model.</p>",
          "rawMarkdown": "Thanks!\n\nI have trained the model using only the new train set.\n\nHowever, since I used pseudo labels to train the model in second stage training.\nIf there is s-shape target that first stage model can determine with high confidence (if it is in top 5,000 positive), it was included in the training data.\n\nAlso, I use soft target for pseudo labels.\nIt is not the value of 0,1, but the value of the probability predicted by the model.",
          "votes": 1
        }
      ]
    },
    {
      "id": 1481706,
      "postDate": "2021-08-19T16:29:10.017Z",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/knjcode\" target=\"_blank\">@knjcode</a> and thanks for sharing, your solution looks very effective and simple, looks you did`t use extra class(s-shape target that only in test data) for training and get very decent score, your best single model (Efficientnet-B4) already got very high score, may I ask which parts/tricks do you think are the key for your solution, or what is the biggest boost for you? And why did you use so many epochs as 60, I guess it would take very long time to train.</p>",
      "rawMarkdown": "Congratulations @knjcode and thanks for sharing, your solution looks very effective and simple, looks you did`t use extra class(s-shape target that only in test data) for training and get very decent score, your best single model (Efficientnet-B4) already got very high score, may I ask which parts/tricks do you think are the key for your solution, or what is the biggest boost for you? And why did you use so many epochs as 60, I guess it would take very long time to train.",
      "replies": [
        {
          "id": 1482302,
          "postDate": "2021-08-20T02:44:54.473Z",
          "content": "<p>Thanks!</p>\n<p>I think the long training epochs was important in my model architecture and training setup this time.</p>\n<p>In the early days when I started working on the competition, I trained models for 20epochs, and most of the time the best score was in the last epoch or just before it.<br>\nSo I increased training epochs to 30 or 60, and the scores improved further.<br>\nEven in training for 30 epochs, the best score was often at or just before the last epoch, but not so in 60 epochs, which is why I finally decided on use 60 epochs.</p>\n<p>I am wondering if there is any difference in the effect of long training epochs of the model with and without Triplet Attention, but since I have not trained long epoch a model without it, I do not know the effect.</p>",
          "rawMarkdown": "Thanks!\n\nI think the long training epochs was important in my model architecture and training setup this time.\n\nIn the early days when I started working on the competition, I trained models for 20epochs, and most of the time the best score was in the last epoch or just before it.\nSo I increased training epochs to 30 or 60, and the scores improved further.\nEven in training for 30 epochs, the best score was often at or just before the last epoch, but not so in 60 epochs, which is why I finally decided on use 60 epochs.\n\nI am wondering if there is any difference in the effect of long training epochs of the model with and without Triplet Attention, but since I have not trained long epoch a model without it, I do not know the effect.",
          "votes": 3
        }
      ]
    },
    {
      "id": 1481070,
      "postDate": "2021-08-19T09:44:58.940Z",
      "content": "<p>Great work! Congratulations!</p>",
      "rawMarkdown": "Great work! Congratulations!"
    },
    {
      "id": 1481025,
      "postDate": "2021-08-19T09:21:00.707Z",
      "content": "<p>Congratulations! when refining the model using pseudo labeling, did you only use the pseudo images or use the pseudo images together with the whole training set?</p>",
      "rawMarkdown": "Congratulations! when refining the model using pseudo labeling, did you only use the pseudo images or use the pseudo images together with the whole training set?",
      "replies": [
        {
          "id": 1481075,
          "postDate": "2021-08-19T09:48:30.910Z",
          "content": "<p>Thanks!</p>\n<p>Refining was also using a 5 fold split on training set, just like the first stage training.<br>\nSo, I added the whole pseudo images to each training fold.</p>",
          "rawMarkdown": "Thanks!\n\nRefining was also using a 5 fold split on training set, just like the first stage training.\nSo, I added the whole pseudo images to each training fold."
        },
        {
          "id": 1481426,
          "postDate": "2021-08-19T13:48:03.617Z",
          "content": "<p>Hi greay work <a href=\"https://www.kaggle.com/knjcode\" target=\"_blank\">@knjcode</a>, where the pseudo-labels on the new test set only, or did you use another piece of data?</p>",
          "rawMarkdown": "Hi greay work @knjcode, where the pseudo-labels on the new test set only, or did you use another piece of data?"
        },
        {
          "id": 1481444,
          "postDate": "2021-08-19T13:58:07.607Z",
          "content": "<p>Hi thanks!</p>\n<p>Pseudo-labels have been added to the new test set only.</p>",
          "rawMarkdown": "Hi thanks!\n\nPseudo-labels have been added to the new test set only."
        },
        {
          "id": 1481553,
          "postDate": "2021-08-19T15:21:18.850Z",
          "content": "<p>Thank you for sharing! Your solution gives me much inspiration!</p>",
          "rawMarkdown": "Thank you for sharing! Your solution gives me much inspiration!"
        }
      ]
    },
    {
      "id": 1480767,
      "postDate": "2021-08-19T07:04:13.577Z",
      "content": "<p>congratulations on the win <a href=\"https://www.kaggle.com/knjcode\" target=\"_blank\">@knjcode</a> , triplet attention seems to be fairly good technique , can you share the code for the same? </p>",
      "rawMarkdown": "congratulations on the win @knjcode , triplet attention seems to be fairly good technique , can you share the code for the same? ",
      "replies": [
        {
          "id": 1480783,
          "postDate": "2021-08-19T07:12:00.987Z",
          "content": "<p>Thanks!</p>\n<p>I used the code in the following link.<br>\n<a href=\"https://github.com/rwightman/pytorch-image-models/blob/499790e117b2c8c1b57780b73d16c28b84db509e/timm/models/layers/triplet.py\" target=\"_blank\">https://github.com/rwightman/pytorch-image-models/blob/499790e117b2c8c1b57780b73d16c28b84db509e/timm/models/layers/triplet.py</a><br>\nI changed the kernel_size of the convolution layer from 7 to 13.</p>\n<p>There is also an official github repository by the authors of the Triplet Attention paper.<br>\n<a href=\"https://github.com/landskape-ai/triplet-attention\" target=\"_blank\">https://github.com/landskape-ai/triplet-attention</a></p>",
          "rawMarkdown": "Thanks!\n\nI used the code in the following link.\nhttps://github.com/rwightman/pytorch-image-models/blob/499790e117b2c8c1b57780b73d16c28b84db509e/timm/models/layers/triplet.py\nI changed the kernel_size of the convolution layer from 7 to 13.\n\nThere is also an official github repository by the authors of the Triplet Attention paper.\nhttps://github.com/landskape-ai/triplet-attention",
          "votes": 1
        },
        {
          "id": 1483351,
          "postDate": "2021-08-20T15:35:57.783Z",
          "content": "<p>Congratulations for landing 3rd place <a href=\"https://www.kaggle.com/knjcode\" target=\"_blank\">@knjcode</a>! Can you share the intuition behind changing the kernel_size from 7 to 13?</p>",
          "rawMarkdown": "Congratulations for landing 3rd place @knjcode! Can you share the intuition behind changing the kernel_size from 7 to 13?"
        },
        {
          "id": 1483553,
          "postDate": "2021-08-20T17:42:04.733Z",
          "content": "<p>very distinct thought.. how can we add this layer. Does Timm model provides any parameter for same <a href=\"https://www.kaggle.com/knjcode\" target=\"_blank\">@knjcode</a> </p>",
          "rawMarkdown": "very distinct thought.. how can we add this layer. Does Timm model provides any parameter for same @knjcode "
        }
      ]
    },
    {
      "id": 1480466,
      "postDate": "2021-08-19T03:23:28.760Z",
      "content": "<p>Congrats!<br>\nCan you share LB with Triplet Attention vs w/o Triplet Attention?</p>",
      "rawMarkdown": "Congrats!\nCan you share LB with Triplet Attention vs w/o Triplet Attention?",
      "replies": [
        {
          "id": 1480520,
          "postDate": "2021-08-19T04:08:18.340Z",
          "content": "<p>Thanks!</p>\n<p>The difference between with and w/o Triplet Attention was confirmed only with resnet34.<br>\nThe CV scores are as follows. (In this case, kernel_size=7)</p>\n<p>with Triplet Attention: 0.88807<br>\nw/o Triplet Attention: 0.88762</p>\n<p>Sorry, I can't check LB because I have already deleted  the model.</p>",
          "rawMarkdown": "Thanks!\n\nThe difference between with and w/o Triplet Attention was confirmed only with resnet34.\nThe CV scores are as follows. (In this case, kernel_size=7)\n\nwith Triplet Attention: 0.88807\nw/o Triplet Attention: 0.88762\n\nSorry, I can't check LB because I have already deleted  the model.",
          "votes": 1
        },
        {
          "id": 1480571,
          "postDate": "2021-08-19T04:49:51.337Z",
          "content": "<p>Thank you for sharing!</p>",
          "rawMarkdown": "Thank you for sharing!"
        },
        {
          "id": 1482287,
          "postDate": "2021-08-20T02:14:11.470Z",
          "content": "<p>thanks for sharing!!<br>\nI think w/o Triplet Attention: 0.88762 is so high for me.<br>\nto achieve 0.88 I used large model( b3_ns pretrained with old data) and large size(768)…<br>\nI would appreciate it if you could tell me more about the experiment using resnet18.</p>",
          "rawMarkdown": "thanks for sharing!!\nI think w/o Triplet Attention: 0.88762 is so high for me.\nto achieve 0.88 I used large model( b3_ns pretrained with old data) and large size(768)...\nI would appreciate it if you could tell me more about the experiment using resnet18."
        },
        {
          "id": 1482315,
          "postDate": "2021-08-20T03:06:08.280Z",
          "content": "<p>Sorry, the model architecture was a typo.</p>\n<p>It was not resnet18, it was resnet34.<br>\n(I also corrected my reply)</p>",
          "rawMarkdown": "Sorry, the model architecture was a typo.\n\nIt was not resnet18, it was resnet34.\n(I also corrected my reply)",
          "votes": 1
        }
      ]
    },
    {
      "id": 1480450,
      "postDate": "2021-08-19T03:15:33.940Z",
      "content": "<p>Congratulations!! How fast did you get when you changed from Swish to ReLU? I want to try next it.</p>",
      "rawMarkdown": "Congratulations!! How fast did you get when you changed from Swish to ReLU? I want to try next it.",
      "replies": [
        {
          "id": 1480552,
          "postDate": "2021-08-19T04:41:16.750Z",
          "content": "<p>Thanks!</p>\n<p>In the comparison with EfficientNet-B4, the training time for 1epoch was reduced from 387 seconds to 347 seconds.</p>",
          "rawMarkdown": "Thanks!\n\nIn the comparison with EfficientNet-B4, the training time for 1epoch was reduced from 387 seconds to 347 seconds.",
          "votes": 2
        }
      ]
    },
    {
      "id": 1480448,
      "postDate": "2021-08-19T03:12:43.803Z",
      "content": "<p>Congratulations, how do you find madgrad performance? (vs Adam, SGD)</p>",
      "rawMarkdown": "Congratulations, how do you find madgrad performance? (vs Adam, SGD)",
      "replies": [
        {
          "id": 1480523,
          "postDate": "2021-08-19T04:13:02.323Z",
          "content": "<p>Thanks!</p>\n<p>I have only used the MADGRAD optimizer in this competion, so I don't know the difference between it and other optimizers.</p>",
          "rawMarkdown": "Thanks!\n\nI have only used the MADGRAD optimizer in this competion, so I don't know the difference between it and other optimizers."
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 1495844,
      "author_name": "Fuco",
      "author_url": "",
      "post_date": "2021-08-29T20:40:06.090000",
      "content": "<p>Congratulations &amp; thanks for sharing the solution.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1484109,
      "author_name": "Bhargav Ram",
      "author_url": "",
      "post_date": "2021-08-21T04:35:13.460000",
      "content": "<p>Congratulations </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1483001,
      "author_name": "Flavjo Xhelollari",
      "author_url": "",
      "post_date": "2021-08-20T11:47:09.177000",
      "content": "<p>Congrats! It is very interesting.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1482729,
      "author_name": "Dilong Wen",
      "author_url": "",
      "post_date": "2021-08-20T08:14:56.077000",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/knjcode\" target=\"_blank\">@knjcode</a> , May i ask how you used the extra class(s-shape target that only in test data) for training,did you added it to training set as extra data or used it as augmentation and set the target to 1 ?  if you manully added to training set, how many samples did you add? </p>",
      "votes": 0,
      "replies": [
        {
          "id": 1482793,
          "author_name": "kenji",
          "author_url": "",
          "post_date": "2021-08-20T09:29:09.163000",
          "content": "<p>Thanks!</p>\n<p>I have trained the model using only the new train set.</p>\n<p>However, since I used pseudo labels to train the model in second stage training.<br>\nIf there is s-shape target that first stage model can determine with high confidence (if it is in top 5,000 positive), it was included in the training data.</p>\n<p>Also, I use soft target for pseudo labels.<br>\nIt is not the value of 0,1, but the value of the probability predicted by the model.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1481706,
      "author_name": "Hao",
      "author_url": "",
      "post_date": "2021-08-19T16:29:10.017000",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/knjcode\" target=\"_blank\">@knjcode</a> and thanks for sharing, your solution looks very effective and simple, looks you did`t use extra class(s-shape target that only in test data) for training and get very decent score, your best single model (Efficientnet-B4) already got very high score, may I ask which parts/tricks do you think are the key for your solution, or what is the biggest boost for you? And why did you use so many epochs as 60, I guess it would take very long time to train.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1482302,
          "author_name": "kenji",
          "author_url": "",
          "post_date": "2021-08-20T02:44:54.473000",
          "content": "<p>Thanks!</p>\n<p>I think the long training epochs was important in my model architecture and training setup this time.</p>\n<p>In the early days when I started working on the competition, I trained models for 20epochs, and most of the time the best score was in the last epoch or just before it.<br>\nSo I increased training epochs to 30 or 60, and the scores improved further.<br>\nEven in training for 30 epochs, the best score was often at or just before the last epoch, but not so in 60 epochs, which is why I finally decided on use 60 epochs.</p>\n<p>I am wondering if there is any difference in the effect of long training epochs of the model with and without Triplet Attention, but since I have not trained long epoch a model without it, I do not know the effect.</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 1481070,
      "author_name": "Qishen Ha",
      "author_url": "",
      "post_date": "2021-08-19T09:44:58.940000",
      "content": "<p>Great work! Congratulations!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1481025,
      "author_name": "hanx",
      "author_url": "",
      "post_date": "2021-08-19T09:21:00.707000",
      "content": "<p>Congratulations! when refining the model using pseudo labeling, did you only use the pseudo images or use the pseudo images together with the whole training set?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1481075,
          "author_name": "kenji",
          "author_url": "",
          "post_date": "2021-08-19T09:48:30.910000",
          "content": "<p>Thanks!</p>\n<p>Refining was also using a 5 fold split on training set, just like the first stage training.<br>\nSo, I added the whole pseudo images to each training fold.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1481426,
          "author_name": "Felipe Bivort Haiek",
          "author_url": "",
          "post_date": "2021-08-19T13:48:03.617000",
          "content": "<p>Hi greay work <a href=\"https://www.kaggle.com/knjcode\" target=\"_blank\">@knjcode</a>, where the pseudo-labels on the new test set only, or did you use another piece of data?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1481444,
          "author_name": "kenji",
          "author_url": "",
          "post_date": "2021-08-19T13:58:07.607000",
          "content": "<p>Hi thanks!</p>\n<p>Pseudo-labels have been added to the new test set only.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1481553,
          "author_name": "hanx",
          "author_url": "",
          "post_date": "2021-08-19T15:21:18.850000",
          "content": "<p>Thank you for sharing! Your solution gives me much inspiration!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1480767,
      "author_name": "Mr_KnowNothing",
      "author_url": "",
      "post_date": "2021-08-19T07:04:13.577000",
      "content": "<p>congratulations on the win <a href=\"https://www.kaggle.com/knjcode\" target=\"_blank\">@knjcode</a> , triplet attention seems to be fairly good technique , can you share the code for the same? </p>",
      "votes": 0,
      "replies": [
        {
          "id": 1480783,
          "author_name": "kenji",
          "author_url": "",
          "post_date": "2021-08-19T07:12:00.987000",
          "content": "<p>Thanks!</p>\n<p>I used the code in the following link.<br>\n<a href=\"https://github.com/rwightman/pytorch-image-models/blob/499790e117b2c8c1b57780b73d16c28b84db509e/timm/models/layers/triplet.py\" target=\"_blank\">https://github.com/rwightman/pytorch-image-models/blob/499790e117b2c8c1b57780b73d16c28b84db509e/timm/models/layers/triplet.py</a><br>\nI changed the kernel_size of the convolution layer from 7 to 13.</p>\n<p>There is also an official github repository by the authors of the Triplet Attention paper.<br>\n<a href=\"https://github.com/landskape-ai/triplet-attention\" target=\"_blank\">https://github.com/landskape-ai/triplet-attention</a></p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1483351,
          "author_name": "_lev_lipinski",
          "author_url": "",
          "post_date": "2021-08-20T15:35:57.783000",
          "content": "<p>Congratulations for landing 3rd place <a href=\"https://www.kaggle.com/knjcode\" target=\"_blank\">@knjcode</a>! Can you share the intuition behind changing the kernel_size from 7 to 13?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1483553,
          "author_name": "Jaideep",
          "author_url": "",
          "post_date": "2021-08-20T17:42:04.733000",
          "content": "<p>very distinct thought.. how can we add this layer. Does Timm model provides any parameter for same <a href=\"https://www.kaggle.com/knjcode\" target=\"_blank\">@knjcode</a> </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1480466,
      "author_name": "dan",
      "author_url": "",
      "post_date": "2021-08-19T03:23:28.760000",
      "content": "<p>Congrats!<br>\nCan you share LB with Triplet Attention vs w/o Triplet Attention?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1480520,
          "author_name": "kenji",
          "author_url": "",
          "post_date": "2021-08-19T04:08:18.340000",
          "content": "<p>Thanks!</p>\n<p>The difference between with and w/o Triplet Attention was confirmed only with resnet34.<br>\nThe CV scores are as follows. (In this case, kernel_size=7)</p>\n<p>with Triplet Attention: 0.88807<br>\nw/o Triplet Attention: 0.88762</p>\n<p>Sorry, I can't check LB because I have already deleted  the model.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1480571,
          "author_name": "dan",
          "author_url": "",
          "post_date": "2021-08-19T04:49:51.337000",
          "content": "<p>Thank you for sharing!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1482287,
          "author_name": "patriot",
          "author_url": "",
          "post_date": "2021-08-20T02:14:11.470000",
          "content": "<p>thanks for sharing!!<br>\nI think w/o Triplet Attention: 0.88762 is so high for me.<br>\nto achieve 0.88 I used large model( b3_ns pretrained with old data) and large size(768)…<br>\nI would appreciate it if you could tell me more about the experiment using resnet18.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1482315,
          "author_name": "kenji",
          "author_url": "",
          "post_date": "2021-08-20T03:06:08.280000",
          "content": "<p>Sorry, the model architecture was a typo.</p>\n<p>It was not resnet18, it was resnet34.<br>\n(I also corrected my reply)</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1480450,
      "author_name": "patriot",
      "author_url": "",
      "post_date": "2021-08-19T03:15:33.940000",
      "content": "<p>Congratulations!! How fast did you get when you changed from Swish to ReLU? I want to try next it.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1480552,
          "author_name": "kenji",
          "author_url": "",
          "post_date": "2021-08-19T04:41:16.750000",
          "content": "<p>Thanks!</p>\n<p>In the comparison with EfficientNet-B4, the training time for 1epoch was reduced from 387 seconds to 347 seconds.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 1480448,
      "author_name": "Gleb",
      "author_url": "",
      "post_date": "2021-08-19T03:12:43.803000",
      "content": "<p>Congratulations, how do you find madgrad performance? (vs Adam, SGD)</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1480523,
          "author_name": "kenji",
          "author_url": "",
          "post_date": "2021-08-19T04:13:02.323000",
          "content": "<p>Thanks!</p>\n<p>I have only used the MADGRAD optimizer in this competion, so I don't know the difference between it and other optimizers.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1480432": "We thank all organizers for this very exciting competition.  \nCongratulations to all who finished the competition and to the winners.\n\n## Summary\n\nBest submission was an emsemble of four EfficientNets with [Convolutional Triplet Attention Module](https://arxiv.org/abs/2010.03045).\n\n**This post was updated on September 1, 2021.**\nThe main changes are as follows.\n1. Added stage0 to train the model for pseudo-label creation.\n2. Already used pseudo-labels in stage1, and fixed that.\n3. Update initial learning rate value.\n\nThe implementation, including training code, has also been released at the following URL.\nhttps://github.com/knjcode/kaggle-seti-2021\n\n\n## Validation and Preprocess\n\n- Use new train data only\n- StratifiedKFold(k=5)\n- Use ON-channels only (819x256) and resize (768x768)\n\n\n## Model Architecture\n\nBackbone -> Triplet Attention -> GeM Pooling -> FC\n\n- Use [EfficientNet](https://arxiv.org/abs/1905.11946) B1, B2, B3 and B4 backbone\n- [Triplet Attention](https://arxiv.org/abs/2010.03045) (increase kernel size from 7 to 13)\n  - Add only one attention layer after the backbone\n  - I used the implementation in [here](https://github.com/rwightman/pytorch-image-models/blob/499790e117b2c8c1b57780b73d16c28b84db509e/timm/models/layers/triplet.py)\n- [GeM Pooling](https://arxiv.org/abs/1711.02512) (p=4)\n- To accelerate the training, replace Swish to ReLU (B3 and B4 only)\n\n\n## Training\n\nThe training process consists of three stages\n\n### Stage 0\n\nThis model will not be used for the final submission.\n\nTraining EfficientNet-B4\n\n- 60epochs DDP AMP\n- Focal Loss (gamma=0.5)\n- [MADGRAD](https://github.com/facebookresearch/madgrad) optimizer\n  - Initial Lr 1e-2 LinearWarmupCosineAnnealingLR (warmup_epochs=1)\n- Mixup (alpha=1.0)\n  - If the target of either data is 1, mixed_target will also be set to 1\n- Data Augmentation\n  - Horizontal and Vertical flip (p=0.5)\n  - ShiftScaleRotate (shift_limit=0.2, scale_limit=0.2, rotate_limit=0)\n  - RandomResizedCrop (scale=[0.8,1.0], ratio=[0.75,1.333])\n\nGenerate Pseudo labels from oof prediction of this model from new test set.  \n(Select 5,000 images each of positive and negative data with high confidence)\n\n### Stage 1\n\nTraining EfficientNet-B1,B2,B3 and B4 by adding pseudo labeled images.\n\nThe training settings are almost the same as stage0.\n\n- 60epochs DDP AMP\n- Focal Loss (gamma=0.5)\n- [MADGRAD](https://github.com/facebookresearch/madgrad) optimizer\n  - Initial Lr 1e-2 or 1e-3 LinearWarmupCosineAnnealingLR (warmup_epochs=5)\n- Mixup (alpha=1.0)\n  - If the target of either data is 1, mixed_target will also be set to 1\n- Data Augmentation\n  - Horizontal and Vertical flip (p=0.5)\n  - ShiftScaleRotate (shift_limit=0.2, scale_limit=0.2, rotate_limit=0)\n  - RandomResizedCrop (scale=[0.8,1.0], ratio=[0.75,1.333])\n\n\n### Stage2\n\nRefine stage1 models and generate submission with TTA\n\nI refined the model with lighter data augmentation than stage1 for 10epochs.\n\n- 10epochs DDP AMP\n- Focal Loss (gamma=0.5)\n- MADGRAD optimizer\n  - Initial Lr 1e-4 LinearWarmupCosineAnnealingLR (warmup_epochs=0)\n- Horizontal and Vertical flip only with Mixup\n- 4xTTA in inference\n  - regular, hflip, vflip, hflip+vflip\n\n\nBest Single Model  (Efficientnet-B4)\n- 5foldCV: 0.9048   PublicLB: 0.80586   PrivateLB: 0.80294\n\nEnsemble of 4 EfficinetNets\n- 5foldCV: 0.9100   PublicLB: 0.80575   PrivateLB: 0.80475\n\n\n## Reference\n\n- [EfficientNet: Rethinking Model Scaling for Convolutional Neural Networks](https://arxiv.org/abs/1905.11946)\n- [Rotate to Attend: Convolutional Triplet Attention Module](https://arxiv.org/abs/2010.03045)\n- [Fine-tuning CNN Image Retrieval with No Human Annotation](https://arxiv.org/abs/1711.02512)\n",
    "1495844": "Congratulations & thanks for sharing the solution.\n\n",
    "1484109": "Congratulations ",
    "1483001": "Congrats! It is very interesting.",
    "1482729": "Congratulations @knjcode , May i ask how you used the extra class(s-shape target that only in test data) for training,did you added it to training set as extra data or used it as augmentation and set the target to 1 ?  if you manully added to training set, how many samples did you add? ",
    "1481706": "Congratulations @knjcode and thanks for sharing, your solution looks very effective and simple, looks you did`t use extra class(s-shape target that only in test data) for training and get very decent score, your best single model (Efficientnet-B4) already got very high score, may I ask which parts/tricks do you think are the key for your solution, or what is the biggest boost for you? And why did you use so many epochs as 60, I guess it would take very long time to train.",
    "1481070": "Great work! Congratulations!",
    "1481025": "Congratulations! when refining the model using pseudo labeling, did you only use the pseudo images or use the pseudo images together with the whole training set?",
    "1480767": "congratulations on the win @knjcode , triplet attention seems to be fairly good technique , can you share the code for the same? ",
    "1480466": "Congrats!\nCan you share LB with Triplet Attention vs w/o Triplet Attention?",
    "1480450": "Congratulations!! How fast did you get when you changed from Swish to ReLU? I want to try next it.",
    "1480448": "Congratulations, how do you find madgrad performance? (vs Adam, SGD)"
  }
}