{
  "id": 176151,
  "title": "5th place solution write-up ",
  "url": "/competitions/landmark-retrieval-2020/discussion/176151",
  "author_name": "NguyenThanhNhan",
  "post_date": "2020-08-20T17:20:45.817000",
  "votes": 37,
  "comment_count": 14,
  "views": 0,
  "content": "<p>I would like to thank my teammate <a href=\"https://www.kaggle.com/aerdem4\" target=\"_blank\">@aerdem4</a> for his work, discussions and ideas throughout this challenge, our first team-up turned out to be quite successful ✊. Congratulations to all the winning teams and solo winners ! Finally, many thanks to the organizers and Kaggle for hosting this interesting competition.<br>\nOur final submission consisted of 4 models from two different architectures (gempool cnn and delg). </p>\n<p><strong>1. Data pre-processing</strong></p>\n<ul>\n<li>Training only on clean train set (81313 classes), validation by computing global average precision metric on a subset of 120000 images (having overlapping classes with train set) randomly sampled from the index set.</li>\n<li>Data augmentations: resize image longer side to [544, 672] then random crop 512x512, RandAugment + Cutout.</li>\n</ul>\n<p><strong>2. Modeling</strong></p>\n<ul>\n<li>Two architectures: CNN (with SEResNeXt50, SEResNeXt101 and ResNeXt101-32x4d as backbones) with CosFace head, and DELG (re-implemented in PyTorch, with SEResNet101 as backbone).</li>\n<li>Generalized mean pooling with frozen p set to 3 was used; bottleneck structure (GEMPool(2048) -&gt; Linear(512) -&gt; BatchNorm1d -&gt; CosFace(81313)) to reduce computation. </li>\n<li>Models were trained either in 10 or 20 epochs with AdamW optimizer and warm-up cosine annealing scheduler.</li>\n<li>Focal loss and label smoothing were better than cross entropy loss</li>\n</ul>\n<p><strong>3. Inference</strong></p>\n<ul>\n<li>Features (512-dim) extracted at scale 1 for each model.</li>\n<li>Concatenate 4 models’ features into a 2048-dim vector.</li>\n<li>Kernel runtime: 8 hour 20 minutes</li>\n</ul>\n<p><strong>4. Public/Private performance</strong></p>\n<table>\n<thead>\n<tr>\n<th>Methods</th>\n<th>Epochs</th>\n<th>Public</th>\n<th>Private</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>resnext101 gem</td>\n<td>20</td>\n<td>0.34596</td>\n<td>0.31024</td>\n</tr>\n<tr>\n<td>seresnext50 gem</td>\n<td>20</td>\n<td>0.3349</td>\n<td>0.29811</td>\n</tr>\n<tr>\n<td>seresnext101 gem</td>\n<td>20</td>\n<td>0.34749</td>\n<td>0.31282</td>\n</tr>\n<tr>\n<td>seresnet101 delg</td>\n<td>10</td>\n<td>0.336</td>\n<td>0.29882</td>\n</tr>\n<tr>\n<td>ensemble</td>\n<td></td>\n<td>0.36644</td>\n<td>0.32878</td>\n</tr>\n</tbody>\n</table>\n<p><strong>5. Ablations on DELG vs GEM</strong><br>\nWe used SEResNeXt50 for all experiments in this section. </p>\n<table>\n<thead>\n<tr>\n<th>Methods</th>\n<th>Epochs</th>\n<th>Public</th>\n<th>Private</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>gem</td>\n<td>10</td>\n<td>0.32163</td>\n<td>0.28434</td>\n</tr>\n<tr>\n<td>gem + self attention block after res5</td>\n<td>15</td>\n<td>0.32006</td>\n<td>0.28187</td>\n</tr>\n<tr>\n<td>gem + online hard neg mining</td>\n<td>15</td>\n<td>0.32357</td>\n<td>0.28687</td>\n</tr>\n<tr>\n<td>gem + focal loss</td>\n<td>20</td>\n<td>0.32634</td>\n<td>0.29054</td>\n</tr>\n<tr>\n<td>delg + focal loss</td>\n<td>20</td>\n<td>0.32928</td>\n<td>0.29263</td>\n</tr>\n<tr>\n<td>gem + focal loss  + cutout</td>\n<td>20</td>\n<td>0.3349</td>\n<td>0.29811</td>\n</tr>\n</tbody>\n</table>\n<p><strong>6. Things that didn't work for us</strong></p>\n<ul>\n<li>Pre-training on v1 dataset hurt.</li>\n<li>Training on concatenated v1 and v2 data.</li>\n<li>Earlier, I trained an EfficientNet B3 and found out that given the same training configs/ epochs, it performed worse than my baseline ResNet50 (which scored 0.299 on public LB). EfficientNets seem to only work well when you train them long enough with hard augmentations. After reading the 1st place solution, I know which experiment I'm gonna run next for the recognition challenge 😁</li>\n</ul>",
  "messages": [
    {
      "id": 979187,
      "postDate": "2020-08-20T17:20:45.817Z",
      "content": "<p>I would like to thank my teammate <a href=\"https://www.kaggle.com/aerdem4\" target=\"_blank\">@aerdem4</a> for his work, discussions and ideas throughout this challenge, our first team-up turned out to be quite successful ✊. Congratulations to all the winning teams and solo winners ! Finally, many thanks to the organizers and Kaggle for hosting this interesting competition.<br>\nOur final submission consisted of 4 models from two different architectures (gempool cnn and delg). </p>\n<p><strong>1. Data pre-processing</strong></p>\n<ul>\n<li>Training only on clean train set (81313 classes), validation by computing global average precision metric on a subset of 120000 images (having overlapping classes with train set) randomly sampled from the index set.</li>\n<li>Data augmentations: resize image longer side to [544, 672] then random crop 512x512, RandAugment + Cutout.</li>\n</ul>\n<p><strong>2. Modeling</strong></p>\n<ul>\n<li>Two architectures: CNN (with SEResNeXt50, SEResNeXt101 and ResNeXt101-32x4d as backbones) with CosFace head, and DELG (re-implemented in PyTorch, with SEResNet101 as backbone).</li>\n<li>Generalized mean pooling with frozen p set to 3 was used; bottleneck structure (GEMPool(2048) -&gt; Linear(512) -&gt; BatchNorm1d -&gt; CosFace(81313)) to reduce computation. </li>\n<li>Models were trained either in 10 or 20 epochs with AdamW optimizer and warm-up cosine annealing scheduler.</li>\n<li>Focal loss and label smoothing were better than cross entropy loss</li>\n</ul>\n<p><strong>3. Inference</strong></p>\n<ul>\n<li>Features (512-dim) extracted at scale 1 for each model.</li>\n<li>Concatenate 4 models’ features into a 2048-dim vector.</li>\n<li>Kernel runtime: 8 hour 20 minutes</li>\n</ul>\n<p><strong>4. Public/Private performance</strong></p>\n<table>\n<thead>\n<tr>\n<th>Methods</th>\n<th>Epochs</th>\n<th>Public</th>\n<th>Private</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>resnext101 gem</td>\n<td>20</td>\n<td>0.34596</td>\n<td>0.31024</td>\n</tr>\n<tr>\n<td>seresnext50 gem</td>\n<td>20</td>\n<td>0.3349</td>\n<td>0.29811</td>\n</tr>\n<tr>\n<td>seresnext101 gem</td>\n<td>20</td>\n<td>0.34749</td>\n<td>0.31282</td>\n</tr>\n<tr>\n<td>seresnet101 delg</td>\n<td>10</td>\n<td>0.336</td>\n<td>0.29882</td>\n</tr>\n<tr>\n<td>ensemble</td>\n<td></td>\n<td>0.36644</td>\n<td>0.32878</td>\n</tr>\n</tbody>\n</table>\n<p><strong>5. Ablations on DELG vs GEM</strong><br>\nWe used SEResNeXt50 for all experiments in this section. </p>\n<table>\n<thead>\n<tr>\n<th>Methods</th>\n<th>Epochs</th>\n<th>Public</th>\n<th>Private</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>gem</td>\n<td>10</td>\n<td>0.32163</td>\n<td>0.28434</td>\n</tr>\n<tr>\n<td>gem + self attention block after res5</td>\n<td>15</td>\n<td>0.32006</td>\n<td>0.28187</td>\n</tr>\n<tr>\n<td>gem + online hard neg mining</td>\n<td>15</td>\n<td>0.32357</td>\n<td>0.28687</td>\n</tr>\n<tr>\n<td>gem + focal loss</td>\n<td>20</td>\n<td>0.32634</td>\n<td>0.29054</td>\n</tr>\n<tr>\n<td>delg + focal loss</td>\n<td>20</td>\n<td>0.32928</td>\n<td>0.29263</td>\n</tr>\n<tr>\n<td>gem + focal loss  + cutout</td>\n<td>20</td>\n<td>0.3349</td>\n<td>0.29811</td>\n</tr>\n</tbody>\n</table>\n<p><strong>6. Things that didn't work for us</strong></p>\n<ul>\n<li>Pre-training on v1 dataset hurt.</li>\n<li>Training on concatenated v1 and v2 data.</li>\n<li>Earlier, I trained an EfficientNet B3 and found out that given the same training configs/ epochs, it performed worse than my baseline ResNet50 (which scored 0.299 on public LB). EfficientNets seem to only work well when you train them long enough with hard augmentations. After reading the 1st place solution, I know which experiment I'm gonna run next for the recognition challenge 😁</li>\n</ul>",
      "rawMarkdown": "I would like to thank my teammate @aerdem4 for his work, discussions and ideas throughout this challenge, our first team-up turned out to be quite successful ✊. Congratulations to all the winning teams and solo winners ! Finally, many thanks to the organizers and Kaggle for hosting this interesting competition.\nOur final submission consisted of 4 models from two different architectures (gempool cnn and delg). \n\n**1. Data pre-processing**\n+ Training only on clean train set (81313 classes), validation by computing global average precision metric on a subset of 120000 images (having overlapping classes with train set) randomly sampled from the index set.\n+ Data augmentations: resize image longer side to [544, 672] then random crop 512x512, RandAugment + Cutout.\n\n**2. Modeling**\n+ Two architectures: CNN (with SEResNeXt50, SEResNeXt101 and ResNeXt101-32x4d as backbones) with CosFace head, and DELG (re-implemented in PyTorch, with SEResNet101 as backbone).\n+ Generalized mean pooling with frozen p set to 3 was used; bottleneck structure (GEMPool(2048) -> Linear(512) -> BatchNorm1d -> CosFace(81313)) to reduce computation. \n+ Models were trained either in 10 or 20 epochs with AdamW optimizer and warm-up cosine annealing scheduler.\n+ Focal loss and label smoothing were better than cross entropy loss\n\n**3. Inference**\n+ Features (512-dim) extracted at scale 1 for each model.\n+ Concatenate 4 models’ features into a 2048-dim vector.\n+ Kernel runtime: 8 hour 20 minutes\n\n**4. Public/Private performance**\n|Methods  |Epochs  |Public  |Private  |\n| --- | --- |\n|resnext101 gem |20  | 0.34596 |0.31024  |\n|seresnext50 gem  |20  |0.3349  |0.29811 |\n|seresnext101 gem  |20  |0.34749  |0.31282  |\n|seresnet101 delg  |10  |0.336  |0.29882  |\n|ensemble  |  |0.36644  |0.32878  |\n\n**5. Ablations on DELG vs GEM**\nWe used SEResNeXt50 for all experiments in this section. \n|Methods  |Epochs  |Public  |Private  |\n| --- | --- |\n|gem |10  | 0.32163 |0.28434  |\n|gem + self attention block after res5  |15  |0.32006  |0.28187 |\n|gem + online hard neg mining  |15  |0.32357  |0.28687  |\n|gem + focal loss  |20  |0.32634  |0.29054  |\n|delg + focal loss  |20  |0.32928  |0.29263  |\n|gem + focal loss  + cutout |20  |0.3349  |0.29811  |\n\n**6. Things that didn't work for us**\n+ Pre-training on v1 dataset hurt.\n+ Training on concatenated v1 and v2 data.\n+ Earlier, I trained an EfficientNet B3 and found out that given the same training configs/ epochs, it performed worse than my baseline ResNet50 (which scored 0.299 on public LB). EfficientNets seem to only work well when you train them long enough with hard augmentations. After reading the 1st place solution, I know which experiment I'm gonna run next for the recognition challenge 😁",
      "votes": 36
    },
    {
      "id": 986718,
      "postDate": "2020-08-26T18:15:17.550Z",
      "content": "<p>Congrats on your well deserved placement <a href=\"https://www.kaggle.com/andy2709\" target=\"_blank\">@andy2709</a> </p>",
      "rawMarkdown": "Congrats on your well deserved placement @andy2709 ",
      "votes": 1
    },
    {
      "id": 982974,
      "postDate": "2020-08-23T22:39:14.227Z",
      "content": "<p>Congratulations for your 5th place. Can you clarify more on what is online hard negative mining? Thanks a ton.  </p>",
      "rawMarkdown": "Congratulations for your 5th place. Can you clarify more on what is online hard negative mining? Thanks a ton.  ",
      "votes": 1,
      "replies": [
        {
          "id": 983695,
          "postDate": "2020-08-24T14:17:13.177Z",
          "content": "<p>You can see an implementation for Online Hard Example Mining (OHEM) here: <a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/128637\" target=\"_blank\">https://www.kaggle.com/c/bengaliai-cv19/discussion/128637</a></p>",
          "rawMarkdown": "You can see an implementation for Online Hard Example Mining (OHEM) here: https://www.kaggle.com/c/bengaliai-cv19/discussion/128637",
          "votes": 3
        },
        {
          "id": 983803,
          "postDate": "2020-08-24T15:48:34.853Z",
          "content": "<p>Amazing thanks. 😄</p>",
          "rawMarkdown": "Amazing thanks. 😄",
          "votes": 1
        }
      ]
    },
    {
      "id": 980428,
      "postDate": "2020-08-21T15:29:20.267Z",
      "content": "<p>Congratz! Well deserved :)</p>",
      "rawMarkdown": "Congratz! Well deserved :)",
      "votes": 1,
      "replies": [
        {
          "id": 980979,
          "postDate": "2020-08-22T04:35:45.063Z",
          "content": "<p>thank you !</p>",
          "rawMarkdown": "thank you !"
        }
      ]
    },
    {
      "id": 979817,
      "postDate": "2020-08-21T06:10:39.943Z",
      "content": "<p>Congrats bro on 5th place. Learn alot from your solution <a href=\"https://www.kaggle.com/andy2709\" target=\"_blank\">@andy2709</a> </p>",
      "rawMarkdown": "Congrats bro on 5th place. Learn alot from your solution @andy2709 ",
      "votes": 1,
      "replies": [
        {
          "id": 980015,
          "postDate": "2020-08-21T09:15:37.127Z",
          "content": "<p>thanks bro <a href=\"https://www.kaggle.com/duykhanh99\" target=\"_blank\">@duykhanh99</a>  👍</p>",
          "rawMarkdown": "thanks bro @duykhanh99  👍"
        }
      ]
    },
    {
      "id": 997876,
      "postDate": "2020-09-04T09:57:28.113Z",
      "content": "<p><a href=\"https://www.kaggle.com/andy2709\" target=\"_blank\">@andy2709</a> Hello! Congratulations for your win! Do you mind disclosing your randaugment &amp; cutout parameters…?</p>",
      "rawMarkdown": "@andy2709 Hello! Congratulations for your win! Do you mind disclosing your randaugment & cutout parameters...?",
      "votes": 2,
      "replies": [
        {
          "id": 998515,
          "postDate": "2020-09-04T19:30:32.333Z",
          "content": "<p>RandAugment with apply_prob=1, num_ops=2, and magnitude randomly chosen from [1, 14]. I only used rotate, translate, shear and flip and removed all the color-based augmentations.<br>\nCutout with apply_prob = 0.5, other params. were default from other implementations (albumentations, torch etc)</p>",
          "rawMarkdown": "RandAugment with apply_prob=1, num_ops=2, and magnitude randomly chosen from [1, 14]. I only used rotate, translate, shear and flip and removed all the color-based augmentations.\nCutout with apply_prob = 0.5, other params. were default from other implementations (albumentations, torch etc)",
          "votes": 1
        },
        {
          "id": 998695,
          "postDate": "2020-09-05T01:20:43.157Z",
          "content": "<p><a href=\"https://www.kaggle.com/andy2709\" target=\"_blank\">@andy2709</a> Great thanks! May I wonder why did you choose to remove color-based augmentations?</p>",
          "rawMarkdown": "@andy2709 Great thanks! May I wonder why did you choose to remove color-based augmentations?"
        }
      ]
    },
    {
      "id": 982904,
      "postDate": "2020-08-23T19:44:25.143Z",
      "content": "<p>Amazing, thank you for the explanation 👍</p>",
      "rawMarkdown": "Amazing, thank you for the explanation 👍",
      "votes": 1
    },
    {
      "id": 1287645,
      "postDate": "2021-04-29T09:05:54.677Z",
      "content": "<p><code>Concatenate 4 models’ features into a 2048-dim vector.</code><br>\nAfter concatenation did you use any normalization? or just used that vector</p>",
      "rawMarkdown": "`Concatenate 4 models’ features into a 2048-dim vector.`\nAfter concatenation did you use any normalization? or just used that vector"
    },
    {
      "id": 986594,
      "postDate": "2020-08-26T16:04:54.013Z",
      "content": "<p>Hi, firstly congrats on your gold. Well Deserved<br>\nCould you elaborate more on your techniques and pipelining of your DELG model?<br>\nThanks in advance!</p>",
      "rawMarkdown": "Hi, firstly congrats on your gold. Well Deserved\nCould you elaborate more on your techniques and pipelining of your DELG model?\nThanks in advance!\n",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 986718,
      "author_name": "Brenda N",
      "author_url": "",
      "post_date": "2020-08-26T18:15:17.550000",
      "content": "<p>Congrats on your well deserved placement <a href=\"https://www.kaggle.com/andy2709\" target=\"_blank\">@andy2709</a> </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 982974,
      "author_name": "torch",
      "author_url": "",
      "post_date": "2020-08-23T22:39:14.227000",
      "content": "<p>Congratulations for your 5th place. Can you clarify more on what is online hard negative mining? Thanks a ton.  </p>",
      "votes": 1,
      "replies": [
        {
          "id": 983695,
          "author_name": "Ahmet Erdem",
          "author_url": "",
          "post_date": "2020-08-24T14:17:13.177000",
          "content": "<p>You can see an implementation for Online Hard Example Mining (OHEM) here: <a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/128637\" target=\"_blank\">https://www.kaggle.com/c/bengaliai-cv19/discussion/128637</a></p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 983803,
          "author_name": "torch",
          "author_url": "",
          "post_date": "2020-08-24T15:48:34.853000",
          "content": "<p>Amazing thanks. 😄</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 980428,
      "author_name": "Eduardo Rocha de Andrade",
      "author_url": "",
      "post_date": "2020-08-21T15:29:20.267000",
      "content": "<p>Congratz! Well deserved :)</p>",
      "votes": 1,
      "replies": [
        {
          "id": 980979,
          "author_name": "NguyenThanhNhan",
          "author_url": "",
          "post_date": "2020-08-22T04:35:45.063000",
          "content": "<p>thank you !</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 979817,
      "author_name": "KhanhVD",
      "author_url": "",
      "post_date": "2020-08-21T06:10:39.943000",
      "content": "<p>Congrats bro on 5th place. Learn alot from your solution <a href=\"https://www.kaggle.com/andy2709\" target=\"_blank\">@andy2709</a> </p>",
      "votes": 1,
      "replies": [
        {
          "id": 980015,
          "author_name": "NguyenThanhNhan",
          "author_url": "",
          "post_date": "2020-08-21T09:15:37.127000",
          "content": "<p>thanks bro <a href=\"https://www.kaggle.com/duykhanh99\" target=\"_blank\">@duykhanh99</a>  👍</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 997876,
      "author_name": "Nicholas Lyu",
      "author_url": "",
      "post_date": "2020-09-04T09:57:28.113000",
      "content": "<p><a href=\"https://www.kaggle.com/andy2709\" target=\"_blank\">@andy2709</a> Hello! Congratulations for your win! Do you mind disclosing your randaugment &amp; cutout parameters…?</p>",
      "votes": 2,
      "replies": [
        {
          "id": 998515,
          "author_name": "NguyenThanhNhan",
          "author_url": "",
          "post_date": "2020-09-04T19:30:32.333000",
          "content": "<p>RandAugment with apply_prob=1, num_ops=2, and magnitude randomly chosen from [1, 14]. I only used rotate, translate, shear and flip and removed all the color-based augmentations.<br>\nCutout with apply_prob = 0.5, other params. were default from other implementations (albumentations, torch etc)</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 998695,
          "author_name": "Nicholas Lyu",
          "author_url": "",
          "post_date": "2020-09-05T01:20:43.157000",
          "content": "<p><a href=\"https://www.kaggle.com/andy2709\" target=\"_blank\">@andy2709</a> Great thanks! May I wonder why did you choose to remove color-based augmentations?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 982904,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-08-23T19:44:25.143000",
      "content": "<p>Amazing, thank you for the explanation 👍</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1287645,
      "author_name": "DeepUnderstanding",
      "author_url": "",
      "post_date": "2021-04-29T09:05:54.677000",
      "content": "<p><code>Concatenate 4 models’ features into a 2048-dim vector.</code><br>\nAfter concatenation did you use any normalization? or just used that vector</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 986594,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-08-26T16:04:54.013000",
      "content": "<p>Hi, firstly congrats on your gold. Well Deserved<br>\nCould you elaborate more on your techniques and pipelining of your DELG model?<br>\nThanks in advance!</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "979187": "I would like to thank my teammate @aerdem4 for his work, discussions and ideas throughout this challenge, our first team-up turned out to be quite successful ✊. Congratulations to all the winning teams and solo winners ! Finally, many thanks to the organizers and Kaggle for hosting this interesting competition.\nOur final submission consisted of 4 models from two different architectures (gempool cnn and delg). \n\n**1. Data pre-processing**\n+ Training only on clean train set (81313 classes), validation by computing global average precision metric on a subset of 120000 images (having overlapping classes with train set) randomly sampled from the index set.\n+ Data augmentations: resize image longer side to [544, 672] then random crop 512x512, RandAugment + Cutout.\n\n**2. Modeling**\n+ Two architectures: CNN (with SEResNeXt50, SEResNeXt101 and ResNeXt101-32x4d as backbones) with CosFace head, and DELG (re-implemented in PyTorch, with SEResNet101 as backbone).\n+ Generalized mean pooling with frozen p set to 3 was used; bottleneck structure (GEMPool(2048) -> Linear(512) -> BatchNorm1d -> CosFace(81313)) to reduce computation. \n+ Models were trained either in 10 or 20 epochs with AdamW optimizer and warm-up cosine annealing scheduler.\n+ Focal loss and label smoothing were better than cross entropy loss\n\n**3. Inference**\n+ Features (512-dim) extracted at scale 1 for each model.\n+ Concatenate 4 models’ features into a 2048-dim vector.\n+ Kernel runtime: 8 hour 20 minutes\n\n**4. Public/Private performance**\n|Methods  |Epochs  |Public  |Private  |\n| --- | --- |\n|resnext101 gem |20  | 0.34596 |0.31024  |\n|seresnext50 gem  |20  |0.3349  |0.29811 |\n|seresnext101 gem  |20  |0.34749  |0.31282  |\n|seresnet101 delg  |10  |0.336  |0.29882  |\n|ensemble  |  |0.36644  |0.32878  |\n\n**5. Ablations on DELG vs GEM**\nWe used SEResNeXt50 for all experiments in this section. \n|Methods  |Epochs  |Public  |Private  |\n| --- | --- |\n|gem |10  | 0.32163 |0.28434  |\n|gem + self attention block after res5  |15  |0.32006  |0.28187 |\n|gem + online hard neg mining  |15  |0.32357  |0.28687  |\n|gem + focal loss  |20  |0.32634  |0.29054  |\n|delg + focal loss  |20  |0.32928  |0.29263  |\n|gem + focal loss  + cutout |20  |0.3349  |0.29811  |\n\n**6. Things that didn't work for us**\n+ Pre-training on v1 dataset hurt.\n+ Training on concatenated v1 and v2 data.\n+ Earlier, I trained an EfficientNet B3 and found out that given the same training configs/ epochs, it performed worse than my baseline ResNet50 (which scored 0.299 on public LB). EfficientNets seem to only work well when you train them long enough with hard augmentations. After reading the 1st place solution, I know which experiment I'm gonna run next for the recognition challenge 😁",
    "986718": "Congrats on your well deserved placement @andy2709 ",
    "982974": "Congratulations for your 5th place. Can you clarify more on what is online hard negative mining? Thanks a ton.  ",
    "980428": "Congratz! Well deserved :)",
    "979817": "Congrats bro on 5th place. Learn alot from your solution @andy2709 ",
    "997876": "@andy2709 Hello! Congratulations for your win! Do you mind disclosing your randaugment & cutout parameters...?",
    "982904": "Amazing, thank you for the explanation 👍",
    "1287645": "`Concatenate 4 models’ features into a 2048-dim vector.`\nAfter concatenation did you use any normalization? or just used that vector",
    "986594": "Hi, firstly congrats on your gold. Well Deserved\nCould you elaborate more on your techniques and pipelining of your DELG model?\nThanks in advance!\n"
  }
}