{
  "id": 176405,
  "title": "8th place solution",
  "url": "/competitions/landmark-retrieval-2020/writeups/tf-tf-tf-8th-place-solution",
  "author_name": "",
  "post_date": "2020-08-21T15:58:03.523Z",
  "votes": 31,
  "comment_count": 11,
  "views": 0,
  "content": "<p>First, thanks to Google and Kaggle for hosting this competition and congratulations to the winners :)</p>\n<p>Since our solution is quite similar to the ones already posted and to DELG itself, I'll keep it simple and brief</p>\n<h3>Dataset</h3>\n<ul>\n<li>GLDv2 clean</li>\n<li>80% training, 20% val split by image</li>\n</ul>\n<h3>Loss</h3>\n<ul>\n<li>ArcFace Layer<ul>\n<li>Margin: 0.3</li>\n<li>Scale: 46</li></ul></li>\n</ul>\n<h3>Models</h3>\n<ul>\n<li>ResNet101</li>\n<li>EfficientNetB5</li>\n<li>GeM pooling<ul>\n<li>p=3 frozen for R101</li>\n<li>p trained for B5</li></ul></li>\n<li>2048d descriptors by applying FC + BN after pool</li>\n</ul>\n<h3>Training</h3>\n<ul>\n<li>Trained until convergence at 512x512</li>\n<li>Fine-tuned for a few epochs at 640x640</li>\n<li>Around 35 epochs for R101 and 20 for B5</li>\n</ul>\n<h3>Inference</h3>\n<ul>\n<li>Multi-scale TTA<ul>\n<li>R101: Resize to (640, 768) squared images </li>\n<li>B5: Resize to 640 squared images + resize to 1024 preserving AR</li>\n<li>2048d descriptors per model by averaging multi-scale predictions followed by l2-normalization</li></ul></li>\n<li>Ensembling by concatenating model's predictions into a 4096d descriptor followed by l2-normalization</li>\n</ul>",
  "messages": [
    {
      "id": "980454",
      "postDate": "08/21/2020 15:54:12",
      "content": "<p>First, thanks to Google and Kaggle for hosting this competition and congratulations to the winners :)</p>\n<p>Since our solution is quite similar to the ones already posted and to DELG itself, I'll keep it simple and brief</p>\n<h3>Dataset</h3>\n<ul>\n<li>GLDv2 clean</li>\n<li>80% training, 20% val split by image</li>\n</ul>\n<h3>Loss</h3>\n<ul>\n<li>ArcFace Layer<ul>\n<li>Margin: 0.3</li>\n<li>Scale: 46</li></ul></li>\n</ul>\n<h3>Models</h3>\n<ul>\n<li>ResNet101</li>\n<li>EfficientNetB5</li>\n<li>GeM pooling<ul>\n<li>p=3 frozen for R101</li>\n<li>p trained for B5</li></ul></li>\n<li>2048d descriptors by applying FC + BN after pool</li>\n</ul>\n<h3>Training</h3>\n<ul>\n<li>Trained until convergence at 512x512</li>\n<li>Fine-tuned for a few epochs at 640x640</li>\n<li>Around 35 epochs for R101 and 20 for B5</li>\n</ul>\n<h3>Inference</h3>\n<ul>\n<li>Multi-scale TTA<ul>\n<li>R101: Resize to (640, 768) squared images </li>\n<li>B5: Resize to 640 squared images + resize to 1024 preserving AR</li>\n<li>2048d descriptors per model by averaging multi-scale predictions followed by l2-normalization</li></ul></li>\n<li>Ensembling by concatenating model's predictions into a 4096d descriptor followed by l2-normalization</li>\n</ul>",
      "rawMarkdown": "First, thanks to Google and Kaggle for hosting this competition and congratulations to the winners :)\n\nSince our solution is quite similar to the ones already posted and to DELG itself, I'll keep it simple and brief\n\n### Dataset\n- GLDv2 clean\n- 80% training, 20% val split by image\n\n### Loss\n- ArcFace Layer\n  - Margin: 0.3\n  - Scale: 46\n\n### Models\n- ResNet101\n- EfficientNetB5\n- GeM pooling\n  - p=3 frozen for R101\n  - p trained for B5\n- 2048d descriptors by applying FC + BN after pool\n\n### Training\n- Trained until convergence at 512x512\n- Fine-tuned for a few epochs at 640x640\n- Around 35 epochs for R101 and 20 for B5\n\n### Inference\n- Multi-scale TTA\n  - R101: Resize to (640, 768) squared images \n  - B5: Resize to 640 squared images + resize to 1024 preserving AR\n  - 2048d descriptors per model by averaging multi-scale predictions followed by l2-normalization\n- Ensembling by concatenating model's predictions into a 4096d descriptor followed by l2-normalization",
      "votes": null
    },
    {
      "id": "981870",
      "postDate": "08/22/2020 19:05:50",
      "content": "<p>Informative. \nThanks for sharing.</p>",
      "rawMarkdown": "Informative. \nThanks for sharing.",
      "votes": null
    },
    {
      "id": "981915",
      "postDate": "08/22/2020 20:48:30",
      "content": "<p>Another gold medal?! Congratulation to you and the team!</p>",
      "rawMarkdown": "Another gold medal?! Congratulation to you and the team!",
      "votes": null
    },
    {
      "id": "981942",
      "postDate": "08/22/2020 21:48:29",
      "content": "<p>Thank you Peter :D</p>",
      "rawMarkdown": "Thank you Peter :D",
      "votes": null
    },
    {
      "id": "981980",
      "postDate": "08/22/2020 23:13:24",
      "content": "<p>Congratulations! How long did it take for your epochs please and what hardware did you use to train?</p>",
      "rawMarkdown": "Congratulations! How long did it take for your epochs please and what hardware did you use to train?",
      "votes": null
    },
    {
      "id": "981990",
      "postDate": "08/22/2020 23:37:33",
      "content": "<p>We only used Kaggle's TPUs. For R101 it was about 1 hour per epoch in 512x512, whereas for B5 it was abour 1h20min</p>",
      "rawMarkdown": "We only used Kaggle's TPUs. For R101 it was about 1 hour per epoch in 512x512, whereas for B5 it was abour 1h20min",
      "votes": null
    },
    {
      "id": "983511",
      "postDate": "08/24/2020 11:20:10",
      "content": "<p>May I know if u r using GPU/TPU? What is the learning rate / schedule? What augmentations did u use? What is the single model single resolution score?</p>",
      "rawMarkdown": "May I know if u r using GPU/TPU? What is the learning rate / schedule? What augmentations did u use? What is the single model single resolution score?",
      "votes": null
    },
    {
      "id": "983613",
      "postDate": "08/24/2020 12:53:12",
      "content": "<p>We used Kaggle's TPU. We used CosineAnnealing schedule (approximately because we did it by hand). No augmentations, just a preprocessing where we would sample an AR and crop it from the original image (similar to DELF's original one). We didn't try single resolution but it should be just a little worse, we got (0.30952/0.34823) for B5.</p>",
      "rawMarkdown": "We used Kaggle's TPU. We used CosineAnnealing schedule (approximately because we did it by hand). No augmentations, just a preprocessing where we would sample an AR and crop it from the original image (similar to DELF's original one). We didn't try single resolution but it should be just a little worse, we got (0.30952/0.34823) for B5.",
      "votes": null
    },
    {
      "id": "983648",
      "postDate": "08/24/2020 13:26:00",
      "content": "<p>Thanks for your reply, if possible, I would like to know also your min and max learning rate. Thank you very much again.</p>",
      "rawMarkdown": "Thanks for your reply, if possible, I would like to know also your min and max learning rate. Thank you very much again.",
      "votes": null
    },
    {
      "id": "983718",
      "postDate": "08/24/2020 14:34:48",
      "content": "<p>1e-3 as base, 8e-4 for the minimum. But we scaled it linearly with batch size since TPUs can fit huge batches.</p>",
      "rawMarkdown": "1e-3 as base, 8e-4 for the minimum. But we scaled it linearly with batch size since TPUs can fit huge batches.",
      "votes": null
    },
    {
      "id": "985895",
      "postDate": "08/26/2020 04:59:39",
      "content": "<p>Congratulations!  Which framework do you use to train efficientNet? Pytorch of TF?</p>",
      "rawMarkdown": "Congratulations!  Which framework do you use to train efficientNet? Pytorch of TF?",
      "votes": null
    },
    {
      "id": "986318",
      "postDate": "08/26/2020 11:26:40",
      "content": "<p>Thanks. I used TF</p>",
      "rawMarkdown": "Thanks. I used TF",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 981915,
      "author_name": "pestipeti",
      "author_url": "",
      "post_date": "08/22/2020 20:48:30",
      "content": "<p>Another gold medal?! Congratulation to you and the team!</p>",
      "votes": null,
      "replies": [
        {
          "id": 981942,
          "author_name": "arc144",
          "author_url": "",
          "post_date": "08/22/2020 21:48:29",
          "content": "<p>Thank you Peter :D</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 981980,
      "author_name": "aroraaman",
      "author_url": "",
      "post_date": "08/22/2020 23:13:24",
      "content": "<p>Congratulations! How long did it take for your epochs please and what hardware did you use to train?</p>",
      "votes": null,
      "replies": [
        {
          "id": 981990,
          "author_name": "arc144",
          "author_url": "",
          "post_date": "08/22/2020 23:37:33",
          "content": "<p>We only used Kaggle's TPUs. For R101 it was about 1 hour per epoch in 512x512, whereas for B5 it was abour 1h20min</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 983511,
      "author_name": "css919",
      "author_url": "",
      "post_date": "08/24/2020 11:20:10",
      "content": "<p>May I know if u r using GPU/TPU? What is the learning rate / schedule? What augmentations did u use? What is the single model single resolution score?</p>",
      "votes": null,
      "replies": [
        {
          "id": 983613,
          "author_name": "arc144",
          "author_url": "",
          "post_date": "08/24/2020 12:53:12",
          "content": "<p>We used Kaggle's TPU. We used CosineAnnealing schedule (approximately because we did it by hand). No augmentations, just a preprocessing where we would sample an AR and crop it from the original image (similar to DELF's original one). We didn't try single resolution but it should be just a little worse, we got (0.30952/0.34823) for B5.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 983648,
          "author_name": "css919",
          "author_url": "",
          "post_date": "08/24/2020 13:26:00",
          "content": "<p>Thanks for your reply, if possible, I would like to know also your min and max learning rate. Thank you very much again.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 983718,
          "author_name": "arc144",
          "author_url": "",
          "post_date": "08/24/2020 14:34:48",
          "content": "<p>1e-3 as base, 8e-4 for the minimum. But we scaled it linearly with batch size since TPUs can fit huge batches.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 985895,
      "author_name": "ll490187880",
      "author_url": "",
      "post_date": "08/26/2020 04:59:39",
      "content": "<p>Congratulations!  Which framework do you use to train efficientNet? Pytorch of TF?</p>",
      "votes": null,
      "replies": [
        {
          "id": 986318,
          "author_name": "arc144",
          "author_url": "",
          "post_date": "08/26/2020 11:26:40",
          "content": "<p>Thanks. I used TF</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 981870,
      "author_name": "raoofnaushad",
      "author_url": "",
      "post_date": "08/22/2020 19:05:50",
      "content": "<p>Informative. \nThanks for sharing.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "980454": "First, thanks to Google and Kaggle for hosting this competition and congratulations to the winners :)\n\nSince our solution is quite similar to the ones already posted and to DELG itself, I'll keep it simple and brief\n\n### Dataset\n- GLDv2 clean\n- 80% training, 20% val split by image\n\n### Loss\n- ArcFace Layer\n  - Margin: 0.3\n  - Scale: 46\n\n### Models\n- ResNet101\n- EfficientNetB5\n- GeM pooling\n  - p=3 frozen for R101\n  - p trained for B5\n- 2048d descriptors by applying FC + BN after pool\n\n### Training\n- Trained until convergence at 512x512\n- Fine-tuned for a few epochs at 640x640\n- Around 35 epochs for R101 and 20 for B5\n\n### Inference\n- Multi-scale TTA\n  - R101: Resize to (640, 768) squared images \n  - B5: Resize to 640 squared images + resize to 1024 preserving AR\n  - 2048d descriptors per model by averaging multi-scale predictions followed by l2-normalization\n- Ensembling by concatenating model's predictions into a 4096d descriptor followed by l2-normalization",
    "981870": "Informative. \nThanks for sharing.",
    "981915": "Another gold medal?! Congratulation to you and the team!",
    "981942": "Thank you Peter :D",
    "981980": "Congratulations! How long did it take for your epochs please and what hardware did you use to train?",
    "981990": "We only used Kaggle's TPUs. For R101 it was about 1 hour per epoch in 512x512, whereas for B5 it was abour 1h20min",
    "983511": "May I know if u r using GPU/TPU? What is the learning rate / schedule? What augmentations did u use? What is the single model single resolution score?",
    "983613": "We used Kaggle's TPU. We used CosineAnnealing schedule (approximately because we did it by hand). No augmentations, just a preprocessing where we would sample an AR and crop it from the original image (similar to DELF's original one). We didn't try single resolution but it should be just a little worse, we got (0.30952/0.34823) for B5.",
    "983648": "Thanks for your reply, if possible, I would like to know also your min and max learning rate. Thank you very much again.",
    "983718": "1e-3 as base, 8e-4 for the minimum. But we scaled it linearly with batch size since TPUs can fit huge batches.",
    "985895": "Congratulations!  Which framework do you use to train efficientNet? Pytorch of TF?",
    "986318": "Thanks. I used TF"
  },
  "source": "meta"
}