{
  "id": 175378,
  "title": "4th place solution and experience sharing",
  "url": "/competitions/landmark-retrieval-2020/writeups/import-tensorflow-as-torch-4th-place-solution-and-",
  "author_name": "",
  "post_date": "2020-08-18T03:50:43.140Z",
  "votes": 58,
  "comment_count": 26,
  "views": 0,
  "content": "<p>Hi all, here's our brief solution writeup:</p>\n<h4>What we have tried and works</h4>\n<ul>\n<li>Augmentation: such as random crop and rotate. Using AutoAugmentation really helps.</li>\n<li>Backbones: Resnest200 and Resnet152 due to the resource limit.</li>\n<li>Pretrain: ImageNet pretrained and Softmax pretrained helps convergence.</li>\n<li>Loss Function: angular based loss function such as ArcFace.</li>\n<li>Label Smoothing</li>\n<li>Cosine learning rate with warmup</li>\n<li>Larger input size. We try 224, 336, 448 and 560. We choose 448 as final input size because of its higher cost performance. Smaller input size causes a drop in score.</li>\n</ul>\n<h4>What we have tried but not works</h4>\n<ul>\n<li>EfficientNet B7: we trained b7 in pytorch and transferred it in TF savedmodel format, but it failed with <strong>Notebook Timeout</strong>.</li>\n<li>Some other large backbones but with lower scores: SEResNext, APolyNet, FishNet, HRNet.</li>\n<li>Some other hyper param in loss function such as larger or smaller margin.</li>\n<li>AdaBN</li>\n<li>DCN</li>\n</ul>\n<h4>What we haven't tried</h4>\n<ul>\n<li>Larger backbones, such as Resnest269. </li>\n<li>It seems that B7 works <a href=\"https://www.kaggle.com/c/landmark-retrieval-2020/discussion/175306\" target=\"_blank\">here</a>. Maybe the way we transfer our models from pytorch to tensorflow causes high time cost in the submission.</li>\n<li>Multi scale input like what baseline model has done.</li>\n<li>EMA</li>\n<li>KD</li>\n</ul>\n<p>Trained models will be upload in a few days. Thanks.</p>",
  "messages": [
    {
      "id": "974779",
      "postDate": "08/18/2020 03:47:21",
      "content": "<p>Hi all, here's our brief solution writeup:</p>\n<h4>What we have tried and works</h4>\n<ul>\n<li>Augmentation: such as random crop and rotate. Using AutoAugmentation really helps.</li>\n<li>Backbones: Resnest200 and Resnet152 due to the resource limit.</li>\n<li>Pretrain: ImageNet pretrained and Softmax pretrained helps convergence.</li>\n<li>Loss Function: angular based loss function such as ArcFace.</li>\n<li>Label Smoothing</li>\n<li>Cosine learning rate with warmup</li>\n<li>Larger input size. We try 224, 336, 448 and 560. We choose 448 as final input size because of its higher cost performance. Smaller input size causes a drop in score.</li>\n</ul>\n<h4>What we have tried but not works</h4>\n<ul>\n<li>EfficientNet B7: we trained b7 in pytorch and transferred it in TF savedmodel format, but it failed with <strong>Notebook Timeout</strong>.</li>\n<li>Some other large backbones but with lower scores: SEResNext, APolyNet, FishNet, HRNet.</li>\n<li>Some other hyper param in loss function such as larger or smaller margin.</li>\n<li>AdaBN</li>\n<li>DCN</li>\n</ul>\n<h4>What we haven't tried</h4>\n<ul>\n<li>Larger backbones, such as Resnest269. </li>\n<li>It seems that B7 works <a href=\"https://www.kaggle.com/c/landmark-retrieval-2020/discussion/175306\" target=\"_blank\">here</a>. Maybe the way we transfer our models from pytorch to tensorflow causes high time cost in the submission.</li>\n<li>Multi scale input like what baseline model has done.</li>\n<li>EMA</li>\n<li>KD</li>\n</ul>\n<p>Trained models will be upload in a few days. Thanks.</p>",
      "rawMarkdown": "Hi all, here's our brief solution writeup:\n\n#### What we have tried and works\n* Augmentation: such as random crop and rotate. Using AutoAugmentation really helps.\n* Backbones: Resnest200 and Resnet152 due to the resource limit.\n* Pretrain: ImageNet pretrained and Softmax pretrained helps convergence.\n* Loss Function: angular based loss function such as ArcFace.\n* Label Smoothing\n* Cosine learning rate with warmup\n* Larger input size. We try 224, 336, 448 and 560. We choose 448 as final input size because of its higher cost performance. Smaller input size causes a drop in score.\n\n#### What we have tried but not works\n* EfficientNet B7: we trained b7 in pytorch and transferred it in TF savedmodel format, but it failed with **Notebook Timeout**.\n* Some other large backbones but with lower scores: SEResNext, APolyNet, FishNet, HRNet.\n* Some other hyper param in loss function such as larger or smaller margin.\n* AdaBN\n* DCN\n\n#### What we haven't tried\n* Larger backbones, such as Resnest269. \n* It seems that B7 works [here](https://www.kaggle.com/c/landmark-retrieval-2020/discussion/175306). Maybe the way we transfer our models from pytorch to tensorflow causes high time cost in the submission.\n* Multi scale input like what baseline model has done.\n* EMA\n* KD\n\nTrained models will be upload in a few days. Thanks.",
      "votes": null
    },
    {
      "id": "974795",
      "postDate": "08/18/2020 03:57:04",
      "content": "<p>Congrats, good job on this competition</p>",
      "rawMarkdown": "Congrats, good job on this competition",
      "votes": null
    },
    {
      "id": "974798",
      "postDate": "08/18/2020 03:57:46",
      "content": "<p>Thank you for your sharing. May I know more about your hardware setup? Do you use GCP or your own computer?</p>",
      "rawMarkdown": "Thank you for your sharing. May I know more about your hardware setup? Do you use GCP or your own computer?",
      "votes": null
    },
    {
      "id": "974806",
      "postDate": "08/18/2020 04:01:09",
      "content": "<p>Thanks. We didn't use GCP.</p>",
      "rawMarkdown": "Thanks. We didn't use GCP.",
      "votes": null
    },
    {
      "id": "974811",
      "postDate": "08/18/2020 04:03:06",
      "content": "<p>Congratulations on your strong finish!</p>\n<p>What was the size of your embeddings, and did you do any dimensionality reduction/whitening?</p>",
      "rawMarkdown": "Congratulations on your strong finish!\n\nWhat was the size of your embeddings, and did you do any dimensionality reduction/whitening?",
      "votes": null
    },
    {
      "id": "974814",
      "postDate": "08/18/2020 04:07:37",
      "content": "<p>The embedding of each model is 512. We concatenated two models' embedding as final embedding with size of 1024. We didn't do whitening.</p>",
      "rawMarkdown": "The embedding of each model is 512. We concatenated two models' embedding as final embedding with size of 1024. We didn't do whitening.",
      "votes": null
    },
    {
      "id": "974831",
      "postDate": "08/18/2020 04:17:55",
      "content": "<p>May I know what kind of GPU cards you used?</p>",
      "rawMarkdown": "May I know what kind of GPU cards you used?",
      "votes": null
    },
    {
      "id": "974838",
      "postDate": "08/18/2020 04:22:59",
      "content": "<p>Congrats, and thanks for sharing! Can you please share more info on training details when you have time?</p>",
      "rawMarkdown": "Congrats, and thanks for sharing! Can you please share more info on training details when you have time?",
      "votes": null
    },
    {
      "id": "974875",
      "postDate": "08/18/2020 04:38:14",
      "content": "<p>Interesting. For me training raw with AutoAugmentation didn't lead to convergence. My new model with AutoAugmentation as a 2nd level training technique is still training……</p>",
      "rawMarkdown": "Interesting. For me training raw with AutoAugmentation didn't lead to convergence. My new model with AutoAugmentation as a 2nd level training technique is still training......",
      "votes": null
    },
    {
      "id": "974884",
      "postDate": "08/18/2020 04:40:42",
      "content": "<p>Thanks, we would like to share anything as we can. If you have any questions, please ask here and we will answer as soon as possible.</p>",
      "rawMarkdown": "Thanks, we would like to share anything as we can. If you have any questions, please ask here and we will answer as soon as possible.",
      "votes": null
    },
    {
      "id": "974887",
      "postDate": "08/18/2020 04:41:19",
      "content": "<p>We use 1080Ti with 11GB memory.</p>",
      "rawMarkdown": "We use 1080Ti with 11GB memory.",
      "votes": null
    },
    {
      "id": "974892",
      "postDate": "08/18/2020 04:43:23",
      "content": "<p>Oh wow with Resnet152, I think the batch size can't be large. What's your typical batch_size for Resnet152? I tried Resnet152 and the batch size can only be single digit…….perhaps something is wrong for my setup.</p>",
      "rawMarkdown": "Oh wow with Resnet152, I think the batch size can't be large. What's your typical batch_size for Resnet152? I tried Resnet152 and the batch size can only be single digit.......perhaps something is wrong for my setup.",
      "votes": null
    },
    {
      "id": "974904",
      "postDate": "08/18/2020 04:48:24",
      "content": "<p>That's right. So we use syncbn to enlarge our batch size. We train on 32GPUs and batch size is 16 for each GPU. Thus finally the total batch size is 512.</p>",
      "rawMarkdown": "That's right. So we use syncbn to enlarge our batch size. We train on 32GPUs and batch size is 16 for each GPU. Thus finally the total batch size is 512.",
      "votes": null
    },
    {
      "id": "974929",
      "postDate": "08/18/2020 05:00:26",
      "content": "<p>Do I understand it correct, that you guys trained on 32 cards?  <br>\nI'm sure there is a little disappointment somwhere for you two ,as you guys missed the podium finish and the prize money, perhaps by tiny fraction. However, congratulations on your excellent performance! </p>\n<p>Can you drop a note on your exeprience with arcface training? How you approached it, what sort of performance gain, you guys observed during training and on leaderboard.</p>",
      "rawMarkdown": "Do I understand it correct, that you guys trained on 32 cards?  \nI'm sure there is a little disappointment somwhere for you two ,as you guys missed the podium finish and the prize money, perhaps by tiny fraction. However, congratulations on your excellent performance! \n\nCan you drop a note on your exeprience with arcface training? How you approached it, what sort of performance gain, you guys observed during training and on leaderboard.",
      "votes": null
    },
    {
      "id": "974930",
      "postDate": "08/18/2020 05:01:09",
      "content": "<p>Having a 32-GPU cluster is impressive!</p>",
      "rawMarkdown": "Having a 32-GPU cluster is impressive!",
      "votes": null
    },
    {
      "id": "974965",
      "postDate": "08/18/2020 05:15:29",
      "content": "<p>In fact, we found that at the end of the training, using smaller batch size as 64 or 128 may get more convergence. This is about your policy of augmentation. Hard augmentation causes severe shaking and leads to bad performance. </p>\n<p>We will share our details with arcface soon. By the way, we use the validation dataset from GLRv1 to check out our performance.  </p>",
      "rawMarkdown": "In fact, we found that at the end of the training, using smaller batch size as 64 or 128 may get more convergence. This is about your policy of augmentation. Hard augmentation causes severe shaking and leads to bad performance. \n\nWe will share our details with arcface soon. By the way, we use the validation dataset from GLRv1 to check out our performance.",
      "votes": null
    },
    {
      "id": "974997",
      "postDate": "08/18/2020 05:34:37",
      "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5338238%2Feabc84a9a475111dc7237a137eefab4e%2F2020-08-18%2013-10-45.png?generation=1597729033891052&amp;alt=media\" alt=\"\"></p>\n<p>Hello! I used Kaggle’s free TPUs for training. Backbone: EfficientNetB6, input_size: 256x256, batch_size: 64x8, I can run 3-4 epochs every 3 hours, and save the weights every time I run. Then load the weights and continue running. The scores are only slightly improved:<br>\n[picture]<br>\nI would like to ask how you deal with the extremely uneven category of landmarks in trainning.</p>",
      "rawMarkdown": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5338238%2Feabc84a9a475111dc7237a137eefab4e%2F2020-08-18%2013-10-45.png?generation=1597729033891052&alt=media)\n\n\nHello! I used Kaggle’s free TPUs for training. Backbone: EfficientNetB6, input_size: 256x256, batch_size: 64x8, I can run 3-4 epochs every 3 hours, and save the weights every time I run. Then load the weights and continue running. The scores are only slightly improved:\n[picture]\nI would like to ask how you deal with the extremely uneven category of landmarks in trainning.",
      "votes": null
    },
    {
      "id": "975007",
      "postDate": "08/18/2020 05:41:38",
      "content": "<p>32 gpu cluster..wowww……</p>",
      "rawMarkdown": "32 gpu cluster..wowww......",
      "votes": null
    },
    {
      "id": "975013",
      "postDate": "08/18/2020 05:44:41",
      "content": "<p>good work. can u pls elaborate more on your data strategy. how you fed images to your model. </p>",
      "rawMarkdown": "good work. can u pls elaborate more on your data strategy. how you fed images to your model.",
      "votes": null
    },
    {
      "id": "975176",
      "postDate": "08/18/2020 07:25:12",
      "content": "<p>We try to remove tails of the dataset (remove classes with fewer than 4 pics), but it doesn't bring any better performance. I think it does no matter even if ignoring this problem (?</p>",
      "rawMarkdown": "We try to remove tails of the dataset (remove classes with fewer than 4 pics), but it doesn't bring any better performance. I think it does no matter even if ignoring this problem (?",
      "votes": null
    },
    {
      "id": "975191",
      "postDate": "08/18/2020 07:32:36",
      "content": "<ol>\n<li>Random crop with scale (0.75-1.25) and ratio (0.8-1.2)</li>\n<li>Add augmentation. We will choose two way from numbers of augmentation (such as shear, ratate, posterize, solarize, equalize, color, contrast, invert, sharpness, etc) and randomly choose the magnitude within the given limitation. More details <a href=\"https://arxiv.org/abs/1805.09501\" target=\"_blank\">here</a>.</li>\n<li>Resize to 448 and normalize.</li>\n</ol>",
      "rawMarkdown": "1. Random crop with scale (0.75-1.25) and ratio (0.8-1.2)\n2. Add augmentation. We will choose two way from numbers of augmentation (such as shear, ratate, posterize, solarize, equalize, color, contrast, invert, sharpness, etc) and randomly choose the magnitude within the given limitation. More details [here](https://arxiv.org/abs/1805.09501).\n3. Resize to 448 and normalize.",
      "votes": null
    },
    {
      "id": "975193",
      "postDate": "08/18/2020 07:33:09",
      "content": "<p>I have also tried to take some measures to deal with this problem, but there is also no benefit. Thank you!</p>",
      "rawMarkdown": "I have also tried to take some measures to deal with this problem, but there is also no benefit. Thank you!",
      "votes": null
    },
    {
      "id": "975283",
      "postDate": "08/18/2020 08:23:41",
      "content": "<p>Congratulations and amazing work. Looking forward to your released code. Thanks.</p>",
      "rawMarkdown": "Congratulations and amazing work. Looking forward to your released code. Thanks.",
      "votes": null
    },
    {
      "id": "975330",
      "postDate": "08/18/2020 08:49:28",
      "content": "<p>Congrats!<br>\nI am interested in how much is the improvement of AutoAugmentation.<br>\nHave you Try external data? (GLDv1 or full-GLDv2)<br>\nThanks!</p>",
      "rawMarkdown": "Congrats!\nI am interested in how much is the improvement of AutoAugmentation.\nHave you Try external data? (GLDv1 or full-GLDv2)\nThanks!",
      "votes": null
    },
    {
      "id": "975717",
      "postDate": "08/18/2020 12:43:35",
      "content": "<p>Having a 32-GPU cluster is impressive!<br>\nnice</p>",
      "rawMarkdown": "Having a 32-GPU cluster is impressive!\nnice",
      "votes": null
    },
    {
      "id": "983855",
      "postDate": "08/24/2020 16:25:26",
      "content": "<p>Congrats! Can you please tell us also your initial, max, and min learning rate, the number of warmup steps and the number of epochs for resnest200?</p>",
      "rawMarkdown": "Congrats! Can you please tell us also your initial, max, and min learning rate, the number of warmup steps and the number of epochs for resnest200?",
      "votes": null
    },
    {
      "id": "999186",
      "postDate": "09/05/2020 13:01:23",
      "content": "<p><a href=\"https://www.kaggle.com/canappeco\" target=\"_blank\">@canappeco</a> Congrats for your position! If you dont mind disclosing, what margin do you use for Arcface loss?</p>",
      "rawMarkdown": "canappeco Congrats for your position! If you dont mind disclosing, what margin do you use for Arcface loss?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 974795,
      "author_name": "bopengiowa",
      "author_url": "",
      "post_date": "08/18/2020 03:57:04",
      "content": "<p>Congrats, good job on this competition</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 974798,
      "author_name": "pukkinming",
      "author_url": "",
      "post_date": "08/18/2020 03:57:46",
      "content": "<p>Thank you for your sharing. May I know more about your hardware setup? Do you use GCP or your own computer?</p>",
      "votes": null,
      "replies": [
        {
          "id": 974806,
          "author_name": "canappeco",
          "author_url": "",
          "post_date": "08/18/2020 04:01:09",
          "content": "<p>Thanks. We didn't use GCP.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 974831,
          "author_name": "pukkinming",
          "author_url": "",
          "post_date": "08/18/2020 04:17:55",
          "content": "<p>May I know what kind of GPU cards you used?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 974887,
          "author_name": "canappeco",
          "author_url": "",
          "post_date": "08/18/2020 04:41:19",
          "content": "<p>We use 1080Ti with 11GB memory.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 974892,
          "author_name": "pukkinming",
          "author_url": "",
          "post_date": "08/18/2020 04:43:23",
          "content": "<p>Oh wow with Resnet152, I think the batch size can't be large. What's your typical batch_size for Resnet152? I tried Resnet152 and the batch size can only be single digit…….perhaps something is wrong for my setup.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 974904,
          "author_name": "canappeco",
          "author_url": "",
          "post_date": "08/18/2020 04:48:24",
          "content": "<p>That's right. So we use syncbn to enlarge our batch size. We train on 32GPUs and batch size is 16 for each GPU. Thus finally the total batch size is 512.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 974929,
          "author_name": "dsnil87",
          "author_url": "",
          "post_date": "08/18/2020 05:00:26",
          "content": "<p>Do I understand it correct, that you guys trained on 32 cards?  <br>\nI'm sure there is a little disappointment somwhere for you two ,as you guys missed the podium finish and the prize money, perhaps by tiny fraction. However, congratulations on your excellent performance! </p>\n<p>Can you drop a note on your exeprience with arcface training? How you approached it, what sort of performance gain, you guys observed during training and on leaderboard.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 974930,
          "author_name": "pukkinming",
          "author_url": "",
          "post_date": "08/18/2020 05:01:09",
          "content": "<p>Having a 32-GPU cluster is impressive!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 974965,
          "author_name": "canappeco",
          "author_url": "",
          "post_date": "08/18/2020 05:15:29",
          "content": "<p>In fact, we found that at the end of the training, using smaller batch size as 64 or 128 may get more convergence. This is about your policy of augmentation. Hard augmentation causes severe shaking and leads to bad performance. </p>\n<p>We will share our details with arcface soon. By the way, we use the validation dataset from GLRv1 to check out our performance.  </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 975007,
          "author_name": "udaygurugubelli",
          "author_url": "",
          "post_date": "08/18/2020 05:41:38",
          "content": "<p>32 gpu cluster..wowww……</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 974811,
      "author_name": "erniechiew",
      "author_url": "",
      "post_date": "08/18/2020 04:03:06",
      "content": "<p>Congratulations on your strong finish!</p>\n<p>What was the size of your embeddings, and did you do any dimensionality reduction/whitening?</p>",
      "votes": null,
      "replies": [
        {
          "id": 974814,
          "author_name": "canappeco",
          "author_url": "",
          "post_date": "08/18/2020 04:07:37",
          "content": "<p>The embedding of each model is 512. We concatenated two models' embedding as final embedding with size of 1024. We didn't do whitening.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 974838,
      "author_name": "chankhavu",
      "author_url": "",
      "post_date": "08/18/2020 04:22:59",
      "content": "<p>Congrats, and thanks for sharing! Can you please share more info on training details when you have time?</p>",
      "votes": null,
      "replies": [
        {
          "id": 974884,
          "author_name": "canappeco",
          "author_url": "",
          "post_date": "08/18/2020 04:40:42",
          "content": "<p>Thanks, we would like to share anything as we can. If you have any questions, please ask here and we will answer as soon as possible.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 974875,
      "author_name": "alexanderliao",
      "author_url": "",
      "post_date": "08/18/2020 04:38:14",
      "content": "<p>Interesting. For me training raw with AutoAugmentation didn't lead to convergence. My new model with AutoAugmentation as a 2nd level training technique is still training……</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 974997,
      "author_name": "yueqiangqin",
      "author_url": "",
      "post_date": "08/18/2020 05:34:37",
      "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5338238%2Feabc84a9a475111dc7237a137eefab4e%2F2020-08-18%2013-10-45.png?generation=1597729033891052&amp;alt=media\" alt=\"\"></p>\n<p>Hello! I used Kaggle’s free TPUs for training. Backbone: EfficientNetB6, input_size: 256x256, batch_size: 64x8, I can run 3-4 epochs every 3 hours, and save the weights every time I run. Then load the weights and continue running. The scores are only slightly improved:<br>\n[picture]<br>\nI would like to ask how you deal with the extremely uneven category of landmarks in trainning.</p>",
      "votes": null,
      "replies": [
        {
          "id": 975176,
          "author_name": "canappeco",
          "author_url": "",
          "post_date": "08/18/2020 07:25:12",
          "content": "<p>We try to remove tails of the dataset (remove classes with fewer than 4 pics), but it doesn't bring any better performance. I think it does no matter even if ignoring this problem (?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 975193,
          "author_name": "yueqiangqin",
          "author_url": "",
          "post_date": "08/18/2020 07:33:09",
          "content": "<p>I have also tried to take some measures to deal with this problem, but there is also no benefit. Thank you!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 975013,
      "author_name": "udaygurugubelli",
      "author_url": "",
      "post_date": "08/18/2020 05:44:41",
      "content": "<p>good work. can u pls elaborate more on your data strategy. how you fed images to your model. </p>",
      "votes": null,
      "replies": [
        {
          "id": 975191,
          "author_name": "canappeco",
          "author_url": "",
          "post_date": "08/18/2020 07:32:36",
          "content": "<ol>\n<li>Random crop with scale (0.75-1.25) and ratio (0.8-1.2)</li>\n<li>Add augmentation. We will choose two way from numbers of augmentation (such as shear, ratate, posterize, solarize, equalize, color, contrast, invert, sharpness, etc) and randomly choose the magnitude within the given limitation. More details <a href=\"https://arxiv.org/abs/1805.09501\" target=\"_blank\">here</a>.</li>\n<li>Resize to 448 and normalize.</li>\n</ol>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 975283,
      "author_name": "rhtsingh",
      "author_url": "",
      "post_date": "08/18/2020 08:23:41",
      "content": "<p>Congratulations and amazing work. Looking forward to your released code. Thanks.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 975330,
      "author_name": "guiguzhixing",
      "author_url": "",
      "post_date": "08/18/2020 08:49:28",
      "content": "<p>Congrats!<br>\nI am interested in how much is the improvement of AutoAugmentation.<br>\nHave you Try external data? (GLDv1 or full-GLDv2)<br>\nThanks!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 975717,
      "author_name": "omsapate",
      "author_url": "",
      "post_date": "08/18/2020 12:43:35",
      "content": "<p>Having a 32-GPU cluster is impressive!<br>\nnice</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 983855,
      "author_name": "css919",
      "author_url": "",
      "post_date": "08/24/2020 16:25:26",
      "content": "<p>Congrats! Can you please tell us also your initial, max, and min learning rate, the number of warmup steps and the number of epochs for resnest200?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 999186,
      "author_name": "roguekk007",
      "author_url": "",
      "post_date": "09/05/2020 13:01:23",
      "content": "<p><a href=\"https://www.kaggle.com/canappeco\" target=\"_blank\">@canappeco</a> Congrats for your position! If you dont mind disclosing, what margin do you use for Arcface loss?</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "974779": "Hi all, here's our brief solution writeup:\n\n#### What we have tried and works\n* Augmentation: such as random crop and rotate. Using AutoAugmentation really helps.\n* Backbones: Resnest200 and Resnet152 due to the resource limit.\n* Pretrain: ImageNet pretrained and Softmax pretrained helps convergence.\n* Loss Function: angular based loss function such as ArcFace.\n* Label Smoothing\n* Cosine learning rate with warmup\n* Larger input size. We try 224, 336, 448 and 560. We choose 448 as final input size because of its higher cost performance. Smaller input size causes a drop in score.\n\n#### What we have tried but not works\n* EfficientNet B7: we trained b7 in pytorch and transferred it in TF savedmodel format, but it failed with **Notebook Timeout**.\n* Some other large backbones but with lower scores: SEResNext, APolyNet, FishNet, HRNet.\n* Some other hyper param in loss function such as larger or smaller margin.\n* AdaBN\n* DCN\n\n#### What we haven't tried\n* Larger backbones, such as Resnest269. \n* It seems that B7 works [here](https://www.kaggle.com/c/landmark-retrieval-2020/discussion/175306). Maybe the way we transfer our models from pytorch to tensorflow causes high time cost in the submission.\n* Multi scale input like what baseline model has done.\n* EMA\n* KD\n\nTrained models will be upload in a few days. Thanks.",
    "974795": "Congrats, good job on this competition",
    "974798": "Thank you for your sharing. May I know more about your hardware setup? Do you use GCP or your own computer?",
    "974806": "Thanks. We didn't use GCP.",
    "974811": "Congratulations on your strong finish!\n\nWhat was the size of your embeddings, and did you do any dimensionality reduction/whitening?",
    "974814": "The embedding of each model is 512. We concatenated two models' embedding as final embedding with size of 1024. We didn't do whitening.",
    "974831": "May I know what kind of GPU cards you used?",
    "974838": "Congrats, and thanks for sharing! Can you please share more info on training details when you have time?",
    "974875": "Interesting. For me training raw with AutoAugmentation didn't lead to convergence. My new model with AutoAugmentation as a 2nd level training technique is still training......",
    "974884": "Thanks, we would like to share anything as we can. If you have any questions, please ask here and we will answer as soon as possible.",
    "974887": "We use 1080Ti with 11GB memory.",
    "974892": "Oh wow with Resnet152, I think the batch size can't be large. What's your typical batch_size for Resnet152? I tried Resnet152 and the batch size can only be single digit.......perhaps something is wrong for my setup.",
    "974904": "That's right. So we use syncbn to enlarge our batch size. We train on 32GPUs and batch size is 16 for each GPU. Thus finally the total batch size is 512.",
    "974929": "Do I understand it correct, that you guys trained on 32 cards?  \nI'm sure there is a little disappointment somwhere for you two ,as you guys missed the podium finish and the prize money, perhaps by tiny fraction. However, congratulations on your excellent performance! \n\nCan you drop a note on your exeprience with arcface training? How you approached it, what sort of performance gain, you guys observed during training and on leaderboard.",
    "974930": "Having a 32-GPU cluster is impressive!",
    "974965": "In fact, we found that at the end of the training, using smaller batch size as 64 or 128 may get more convergence. This is about your policy of augmentation. Hard augmentation causes severe shaking and leads to bad performance. \n\nWe will share our details with arcface soon. By the way, we use the validation dataset from GLRv1 to check out our performance.",
    "974997": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5338238%2Feabc84a9a475111dc7237a137eefab4e%2F2020-08-18%2013-10-45.png?generation=1597729033891052&alt=media)\n\n\nHello! I used Kaggle’s free TPUs for training. Backbone: EfficientNetB6, input_size: 256x256, batch_size: 64x8, I can run 3-4 epochs every 3 hours, and save the weights every time I run. Then load the weights and continue running. The scores are only slightly improved:\n[picture]\nI would like to ask how you deal with the extremely uneven category of landmarks in trainning.",
    "975007": "32 gpu cluster..wowww......",
    "975013": "good work. can u pls elaborate more on your data strategy. how you fed images to your model.",
    "975176": "We try to remove tails of the dataset (remove classes with fewer than 4 pics), but it doesn't bring any better performance. I think it does no matter even if ignoring this problem (?",
    "975191": "1. Random crop with scale (0.75-1.25) and ratio (0.8-1.2)\n2. Add augmentation. We will choose two way from numbers of augmentation (such as shear, ratate, posterize, solarize, equalize, color, contrast, invert, sharpness, etc) and randomly choose the magnitude within the given limitation. More details [here](https://arxiv.org/abs/1805.09501).\n3. Resize to 448 and normalize.",
    "975193": "I have also tried to take some measures to deal with this problem, but there is also no benefit. Thank you!",
    "975283": "Congratulations and amazing work. Looking forward to your released code. Thanks.",
    "975330": "Congrats!\nI am interested in how much is the improvement of AutoAugmentation.\nHave you Try external data? (GLDv1 or full-GLDv2)\nThanks!",
    "975717": "Having a 32-GPU cluster is impressive!\nnice",
    "983855": "Congrats! Can you please tell us also your initial, max, and min learning rate, the number of warmup steps and the number of epochs for resnest200?",
    "999186": "canappeco Congrats for your position! If you dont mind disclosing, what margin do you use for Arcface loss?"
  },
  "source": "meta"
}