{
  "id": 277098,
  "title": "1st place solution",
  "url": "/competitions/landmark-recognition-2021/discussion/277098",
  "author_name": "Dieter",
  "post_date": "2021-10-07T21:27:15.163000",
  "votes": 184,
  "comment_count": 34,
  "views": 0,
  "content": "<p>First, let me thank kaggle staff and google team for organizing the 2021 landmark competitions. I am still overwhelmed and shocked by my result, which resulted in becoming #1 on the overall kaggle leaderboard, an ultimate goal for many kagglers. My solutions are similar in terms of models and training routine for the retrieval and the recognition task, but slightly different in details, especially in postprocessing, i.e. in re-ranking procedure. I try to give attribution to people, whose smart ideas I levered from last year's solutions. For recognition track I re-used several ideas from 2nd place solution of GLR2020 by <a href=\"https://www.kaggle.com/bestfitting\" target=\"_blank\">@bestfitting</a>, 3rd place solution of team All Data Are Ext ( <a href=\"https://www.kaggle.com/boliu0\" target=\"_blank\">@boliu0</a> <a href=\"https://www.kaggle.com/haqishen\" target=\"_blank\">@haqishen</a> <a href=\"https://www.kaggle.com/garybios\" target=\"_blank\">@garybios</a> <a href=\"https://www.kaggle.com/alexanderliao\" target=\"_blank\">@alexanderliao</a> ) Please comment if I forget someone.</p>\n<h2>TLDR</h2>\n<p>My solution is an ensemble of DOLG models [1] and Hybrid-Swin-Transformers, which are both especially strong architectures for integrating local and global descriptors of a landmark into a single feature vector in a single-stage manner. I trained the models using a 3-step approach where I enlarged the training dataset and image size in each step. For postprocessing, which mainly handles non-landmark identification, I used the same method as in last year's 1st place recognition solution.</p>\n<h2>Preamble</h2>\n<p>Before going into details I want to address a few circumstances leading me to the solution. Firstly, compared to last year, this year's competition had a significantly more tense timeline, which forced me to skip following some ideas I had and concentrate on the most promising. That especially led to the choice to not consider spatial verification with local descriptors at all, as it has (at least in my opinion) a bad cost-benefit trade-off, especially in large scale image recognition, with limited kernel runtime. Secondly, since top teams of last year's iterations published their solutions in detail, accessible to everybody, the bar to outperform here was quite high. To be consistent in naming I use the following for subsets of gldv2</p>\n<p><strong>gldv2</strong> available under [2] (5 mio images, 200k classes)<br>\n<strong>gldv2c</strong> (cleaned subset of gldv2 / this years training data) (1.6 million training images, 81k classes)<br>\n<strong>gldv2x</strong>, uncleaned subset of gldv2 restricted to only have the same 81k classes as gldv2c (3.2 mio images)</p>\n<p>Furthermore I use <strong>GLRec</strong> for past google landmark recognition competitions and <strong>GLRet</strong> for retrieval. </p>\n<h2>Cross-validation</h2>\n<p>I used the same cross-validation scheme as last year, namely the 2019 labeled test set, and had a good correlation between Public LB and my local validation.</p>\n<h2>Training routine</h2>\n<p>My training routine follows train dataset choice from 2nd place GLRec 2020 but with some improvements using ideas from <em>All Data Are Ext</em><br>\nAs a first step I train models on small image size (e.g. 224x224) on the clean gldv2c for around 10 epochs. Then I use medium large image size (e.g. 512x512) and train for a long time on the more noisy gldv2x (30-40 epochs). Finally I finetune on large image size (e.g. 768x768) also on gldv2x for a few epochs. I used Adam optimizer and cosine annealing schedule in every step. I used same augmentation as <em>All Data Are Ext</em>, which proved to be very efficient:</p>\n<pre><code>cfg.train_aug = A.Compose([\n        A.HorizontalFlip(p=0.5),\n        A.ImageCompression(quality_lower=99, quality_upper=100),\n        A.ShiftScaleRotate(shift_limit=0.2, scale_limit=0.2, rotate_limit=10, border_mode=0, p=0.7),\n        A.Resize(image_size, image_size),\n        A.Cutout(max_h_size=int(image_size * 0.4), max_w_size=int(image_size * 0.4), num_holes=1, p=0.5),\n    ])\n\ncfg.val_aug = A.Compose([A.Resize(image_size, image_size),])\n</code></pre>\n<h2>Model Architectures</h2>\n<p>For me this is the most fun part in every competition, to design and explore model architectures. I like the idea of models solving multiple parts of the problem at hand in end2end fashion. In landmark recognition it is crucial to leverage both, global and local image information. While historically this was solved in a two-stage fashion by training a global descriptor for preliminary recognition on the one hand and extracting local descriptors for post reranking using spatial verification on the other hand. Recently attempts have been made to train local and global descriptors simultaneously using a single model (DELF/ DELG). But that still uses a 2nd stage spatial verification. <br>\nI implemented two architectures that combine local and global descriptor into a single image embedding. Both use an effientnet encoder and an sub-arcface head with dynamic margins which was shown by team <em>All Data Are Ext</em> in 2020 to outperform classic arcface.</p>\n<h3>DOLG</h3>\n<p>The very recently released DOLG paper, goes on step further claiming to train local and global descriptors also simultaneously but additionally fusing them within the model into a single descriptor. Unfortunately/ fortunately it was so new that no code was released so far and I needed to implement deep orthogonal fusion by myself. Luckily the paper gives enough details and I was able to implement it quickly. The following shows the final architecture:</p>\n<p><img src=\"https://i.imgur.com/9FiSEet.png\" alt=\"DOLG-EffiencentNet\"></p>\n<h3>Hybrid Swin Transformer</h3>\n<p>Another architecture I explored was a Hybrid-Transformer, half a CNN encoder with a transformer ending. The idea behind is that transformers are like graph neural nets on image patches connecting distant local features with each other. A property CNNs are lacking, but which are very important for landmark recognition/ retrieval. However, pretrained image transformers are mostly done on 224 or 384 image size, and I wanted to go bigger. Also the encoding of each patch is of a higher quality when using a sophisticated CNN for the patch embedding. Luckily the awesome timm repository gives you the tool set to glue together a CNN and an image transformer and I ended with an architecture like shown below. In particular I used Swin transformer [3], as I think the moving window approach captures more diverse local descriptors and EfficientNet for encoding.</p>\n<p><img src=\"https://i.imgur.com/2iXNuBA.png\" alt=\"Hybrid-Swin-Transformer\"></p>\n<p>However, for training several problems arise: Firstly the pretrained weights of image transformer and CNN are trained separately, hence not fitting very well together, resulting in NaNs while training. Secondly, since image transformer require a fixed input size it's a bit more difficult to follow a training routine with increasing image size in each step. I found that the following procedure fixes the issues and gives a good final model, and especially step 3 is important to make the Hybrid-Swin-Transformer working</p>\n<ol>\n<li>train only image transformer on 224x224 size</li>\n<li>Exchange original patch embedding module with block 0,1,2 from Effnet encoder, </li>\n<li>freeze image transformer and sub-arcface head and train for 1 epoch on 448x448, to let the effnet encoder “adjust”</li>\n<li>unfreeze image transformer and head and train for 30-40 epochs</li>\n<li>add block 3 from effnet encoder and finetune a few epochs on 896x896</li>\n</ol>\n<p>For two of my models I used the “stride trick”, i.e. putting a stride of (1,1) in the first conv layer to effectively double the used resolution instead of double the actual image size, which I think is better for small original images.<br>\nI also retrained a few models from 2nd and 3rd place solutions from 2020 recognition competition following their provided git repositories and instructions, and evaluated possible ensembles on my local cross validation. Models form All Data Ext team where quite low correlated to mine and added a nice diversification benefit</p>\n<p>In my final ensemble I had the following models:</p>\n<ul>\n<li>DOLG-Efficentnet b5 (768x768)</li>\n<li>DOLG-Efficentnet b6 (768x768)</li>\n<li>DOLG-Efficentnet b7 (448x448) stride 1</li>\n<li>Hybrid SwinBase224-Efficentnet b5 (448x448) stride 1</li>\n<li>Hybrid SwinBase224-Efficentnet b3 (896x896)</li>\n<li>Hybrid SwinBase384-Efficentnet b6 (384,384) stride 1</li>\n<li>All Data Ext 2020 Efficientnet b6 (512x512) </li>\n<li>All Data Ext 2020 EfficientNet b3 (768x768)</li>\n</ul>\n<p>Post-processing<br>\nI literally used the same approach (developed by <a href=\"https://www.kaggle.com/philippsinger\" target=\"_blank\">@philippsinger</a>) as in our last year solution. you can find it here:</p>\n<p><a href=\"https://www.kaggle.com/c/landmark-recognition-2020/discussion/187821\" target=\"_blank\">https://www.kaggle.com/c/landmark-recognition-2020/discussion/187821</a></p>\n<p>For more details please refer to the paper on arxiv. Thank you for reading.</p>\n<p>[1] <a href=\"https://arxiv.org/abs/2108.02927\" target=\"_blank\">https://arxiv.org/abs/2108.02927</a> (DOLG)<br>\n[2] <a href=\"https://github.com/cvdfoundation/google-landmark\" target=\"_blank\">https://github.com/cvdfoundation/google-landmark</a><br>\n[3] <a href=\"https://arxiv.org/abs/2103.14030\" target=\"_blank\">https://arxiv.org/abs/2103.14030</a> (Swin)</p>\n<p>Code: <a href=\"https://github.com/ChristofHenkel/kaggle-landmark-2021-1st-place\" target=\"_blank\">https://github.com/ChristofHenkel/kaggle-landmark-2021-1st-place</a> (in progress) <br>\nPaper: submitted to arxiv, waiting for acception. In the meanwhile I uploaded to <a href=\"https://github.com/ChristofHenkel/kaggle-landmark-2021-1st-place/blob/main/GLR_2021_1st_place.pdf\" target=\"_blank\">https://github.com/ChristofHenkel/kaggle-landmark-2021-1st-place/blob/main/GLR_2021_1st_place.pdf</a></p>",
  "messages": [
    {
      "id": 1537908,
      "postDate": "2021-10-07T21:27:15.163Z",
      "content": "<p>First, let me thank kaggle staff and google team for organizing the 2021 landmark competitions. I am still overwhelmed and shocked by my result, which resulted in becoming #1 on the overall kaggle leaderboard, an ultimate goal for many kagglers. My solutions are similar in terms of models and training routine for the retrieval and the recognition task, but slightly different in details, especially in postprocessing, i.e. in re-ranking procedure. I try to give attribution to people, whose smart ideas I levered from last year's solutions. For recognition track I re-used several ideas from 2nd place solution of GLR2020 by <a href=\"https://www.kaggle.com/bestfitting\" target=\"_blank\">@bestfitting</a>, 3rd place solution of team All Data Are Ext ( <a href=\"https://www.kaggle.com/boliu0\" target=\"_blank\">@boliu0</a> <a href=\"https://www.kaggle.com/haqishen\" target=\"_blank\">@haqishen</a> <a href=\"https://www.kaggle.com/garybios\" target=\"_blank\">@garybios</a> <a href=\"https://www.kaggle.com/alexanderliao\" target=\"_blank\">@alexanderliao</a> ) Please comment if I forget someone.</p>\n<h2>TLDR</h2>\n<p>My solution is an ensemble of DOLG models [1] and Hybrid-Swin-Transformers, which are both especially strong architectures for integrating local and global descriptors of a landmark into a single feature vector in a single-stage manner. I trained the models using a 3-step approach where I enlarged the training dataset and image size in each step. For postprocessing, which mainly handles non-landmark identification, I used the same method as in last year's 1st place recognition solution.</p>\n<h2>Preamble</h2>\n<p>Before going into details I want to address a few circumstances leading me to the solution. Firstly, compared to last year, this year's competition had a significantly more tense timeline, which forced me to skip following some ideas I had and concentrate on the most promising. That especially led to the choice to not consider spatial verification with local descriptors at all, as it has (at least in my opinion) a bad cost-benefit trade-off, especially in large scale image recognition, with limited kernel runtime. Secondly, since top teams of last year's iterations published their solutions in detail, accessible to everybody, the bar to outperform here was quite high. To be consistent in naming I use the following for subsets of gldv2</p>\n<p><strong>gldv2</strong> available under [2] (5 mio images, 200k classes)<br>\n<strong>gldv2c</strong> (cleaned subset of gldv2 / this years training data) (1.6 million training images, 81k classes)<br>\n<strong>gldv2x</strong>, uncleaned subset of gldv2 restricted to only have the same 81k classes as gldv2c (3.2 mio images)</p>\n<p>Furthermore I use <strong>GLRec</strong> for past google landmark recognition competitions and <strong>GLRet</strong> for retrieval. </p>\n<h2>Cross-validation</h2>\n<p>I used the same cross-validation scheme as last year, namely the 2019 labeled test set, and had a good correlation between Public LB and my local validation.</p>\n<h2>Training routine</h2>\n<p>My training routine follows train dataset choice from 2nd place GLRec 2020 but with some improvements using ideas from <em>All Data Are Ext</em><br>\nAs a first step I train models on small image size (e.g. 224x224) on the clean gldv2c for around 10 epochs. Then I use medium large image size (e.g. 512x512) and train for a long time on the more noisy gldv2x (30-40 epochs). Finally I finetune on large image size (e.g. 768x768) also on gldv2x for a few epochs. I used Adam optimizer and cosine annealing schedule in every step. I used same augmentation as <em>All Data Are Ext</em>, which proved to be very efficient:</p>\n<pre><code>cfg.train_aug = A.Compose([\n        A.HorizontalFlip(p=0.5),\n        A.ImageCompression(quality_lower=99, quality_upper=100),\n        A.ShiftScaleRotate(shift_limit=0.2, scale_limit=0.2, rotate_limit=10, border_mode=0, p=0.7),\n        A.Resize(image_size, image_size),\n        A.Cutout(max_h_size=int(image_size * 0.4), max_w_size=int(image_size * 0.4), num_holes=1, p=0.5),\n    ])\n\ncfg.val_aug = A.Compose([A.Resize(image_size, image_size),])\n</code></pre>\n<h2>Model Architectures</h2>\n<p>For me this is the most fun part in every competition, to design and explore model architectures. I like the idea of models solving multiple parts of the problem at hand in end2end fashion. In landmark recognition it is crucial to leverage both, global and local image information. While historically this was solved in a two-stage fashion by training a global descriptor for preliminary recognition on the one hand and extracting local descriptors for post reranking using spatial verification on the other hand. Recently attempts have been made to train local and global descriptors simultaneously using a single model (DELF/ DELG). But that still uses a 2nd stage spatial verification. <br>\nI implemented two architectures that combine local and global descriptor into a single image embedding. Both use an effientnet encoder and an sub-arcface head with dynamic margins which was shown by team <em>All Data Are Ext</em> in 2020 to outperform classic arcface.</p>\n<h3>DOLG</h3>\n<p>The very recently released DOLG paper, goes on step further claiming to train local and global descriptors also simultaneously but additionally fusing them within the model into a single descriptor. Unfortunately/ fortunately it was so new that no code was released so far and I needed to implement deep orthogonal fusion by myself. Luckily the paper gives enough details and I was able to implement it quickly. The following shows the final architecture:</p>\n<p><img src=\"https://i.imgur.com/9FiSEet.png\" alt=\"DOLG-EffiencentNet\"></p>\n<h3>Hybrid Swin Transformer</h3>\n<p>Another architecture I explored was a Hybrid-Transformer, half a CNN encoder with a transformer ending. The idea behind is that transformers are like graph neural nets on image patches connecting distant local features with each other. A property CNNs are lacking, but which are very important for landmark recognition/ retrieval. However, pretrained image transformers are mostly done on 224 or 384 image size, and I wanted to go bigger. Also the encoding of each patch is of a higher quality when using a sophisticated CNN for the patch embedding. Luckily the awesome timm repository gives you the tool set to glue together a CNN and an image transformer and I ended with an architecture like shown below. In particular I used Swin transformer [3], as I think the moving window approach captures more diverse local descriptors and EfficientNet for encoding.</p>\n<p><img src=\"https://i.imgur.com/2iXNuBA.png\" alt=\"Hybrid-Swin-Transformer\"></p>\n<p>However, for training several problems arise: Firstly the pretrained weights of image transformer and CNN are trained separately, hence not fitting very well together, resulting in NaNs while training. Secondly, since image transformer require a fixed input size it's a bit more difficult to follow a training routine with increasing image size in each step. I found that the following procedure fixes the issues and gives a good final model, and especially step 3 is important to make the Hybrid-Swin-Transformer working</p>\n<ol>\n<li>train only image transformer on 224x224 size</li>\n<li>Exchange original patch embedding module with block 0,1,2 from Effnet encoder, </li>\n<li>freeze image transformer and sub-arcface head and train for 1 epoch on 448x448, to let the effnet encoder “adjust”</li>\n<li>unfreeze image transformer and head and train for 30-40 epochs</li>\n<li>add block 3 from effnet encoder and finetune a few epochs on 896x896</li>\n</ol>\n<p>For two of my models I used the “stride trick”, i.e. putting a stride of (1,1) in the first conv layer to effectively double the used resolution instead of double the actual image size, which I think is better for small original images.<br>\nI also retrained a few models from 2nd and 3rd place solutions from 2020 recognition competition following their provided git repositories and instructions, and evaluated possible ensembles on my local cross validation. Models form All Data Ext team where quite low correlated to mine and added a nice diversification benefit</p>\n<p>In my final ensemble I had the following models:</p>\n<ul>\n<li>DOLG-Efficentnet b5 (768x768)</li>\n<li>DOLG-Efficentnet b6 (768x768)</li>\n<li>DOLG-Efficentnet b7 (448x448) stride 1</li>\n<li>Hybrid SwinBase224-Efficentnet b5 (448x448) stride 1</li>\n<li>Hybrid SwinBase224-Efficentnet b3 (896x896)</li>\n<li>Hybrid SwinBase384-Efficentnet b6 (384,384) stride 1</li>\n<li>All Data Ext 2020 Efficientnet b6 (512x512) </li>\n<li>All Data Ext 2020 EfficientNet b3 (768x768)</li>\n</ul>\n<p>Post-processing<br>\nI literally used the same approach (developed by <a href=\"https://www.kaggle.com/philippsinger\" target=\"_blank\">@philippsinger</a>) as in our last year solution. you can find it here:</p>\n<p><a href=\"https://www.kaggle.com/c/landmark-recognition-2020/discussion/187821\" target=\"_blank\">https://www.kaggle.com/c/landmark-recognition-2020/discussion/187821</a></p>\n<p>For more details please refer to the paper on arxiv. Thank you for reading.</p>\n<p>[1] <a href=\"https://arxiv.org/abs/2108.02927\" target=\"_blank\">https://arxiv.org/abs/2108.02927</a> (DOLG)<br>\n[2] <a href=\"https://github.com/cvdfoundation/google-landmark\" target=\"_blank\">https://github.com/cvdfoundation/google-landmark</a><br>\n[3] <a href=\"https://arxiv.org/abs/2103.14030\" target=\"_blank\">https://arxiv.org/abs/2103.14030</a> (Swin)</p>\n<p>Code: <a href=\"https://github.com/ChristofHenkel/kaggle-landmark-2021-1st-place\" target=\"_blank\">https://github.com/ChristofHenkel/kaggle-landmark-2021-1st-place</a> (in progress) <br>\nPaper: submitted to arxiv, waiting for acception. In the meanwhile I uploaded to <a href=\"https://github.com/ChristofHenkel/kaggle-landmark-2021-1st-place/blob/main/GLR_2021_1st_place.pdf\" target=\"_blank\">https://github.com/ChristofHenkel/kaggle-landmark-2021-1st-place/blob/main/GLR_2021_1st_place.pdf</a></p>",
      "rawMarkdown": "First, let me thank kaggle staff and google team for organizing the 2021 landmark competitions. I am still overwhelmed and shocked by my result, which resulted in becoming #1 on the overall kaggle leaderboard, an ultimate goal for many kagglers. My solutions are similar in terms of models and training routine for the retrieval and the recognition task, but slightly different in details, especially in postprocessing, i.e. in re-ranking procedure. I try to give attribution to people, whose smart ideas I levered from last year's solutions. For recognition track I re-used several ideas from 2nd place solution of GLR2020 by @bestfitting, 3rd place solution of team All Data Are Ext ( @boliu0 @haqishen @garybios @alexanderliao ) Please comment if I forget someone.\n\n## TLDR\nMy solution is an ensemble of DOLG models [1] and Hybrid-Swin-Transformers, which are both especially strong architectures for integrating local and global descriptors of a landmark into a single feature vector in a single-stage manner. I trained the models using a 3-step approach where I enlarged the training dataset and image size in each step. For postprocessing, which mainly handles non-landmark identification, I used the same method as in last year's 1st place recognition solution.\n\n## Preamble\nBefore going into details I want to address a few circumstances leading me to the solution. Firstly, compared to last year, this year's competition had a significantly more tense timeline, which forced me to skip following some ideas I had and concentrate on the most promising. That especially led to the choice to not consider spatial verification with local descriptors at all, as it has (at least in my opinion) a bad cost-benefit trade-off, especially in large scale image recognition, with limited kernel runtime. Secondly, since top teams of last year's iterations published their solutions in detail, accessible to everybody, the bar to outperform here was quite high. To be consistent in naming I use the following for subsets of gldv2\n\n**gldv2** available under [2] (5 mio images, 200k classes)\n**gldv2c** (cleaned subset of gldv2 / this years training data) (1.6 million training images, 81k classes)\n**gldv2x**, uncleaned subset of gldv2 restricted to only have the same 81k classes as gldv2c (3.2 mio images)\n\nFurthermore I use **GLRec** for past google landmark recognition competitions and **GLRet** for retrieval. \n\n## Cross-validation\nI used the same cross-validation scheme as last year, namely the 2019 labeled test set, and had a good correlation between Public LB and my local validation.\n\n## Training routine\nMy training routine follows train dataset choice from 2nd place GLRec 2020 but with some improvements using ideas from *All Data Are Ext*\nAs a first step I train models on small image size (e.g. 224x224) on the clean gldv2c for around 10 epochs. Then I use medium large image size (e.g. 512x512) and train for a long time on the more noisy gldv2x (30-40 epochs). Finally I finetune on large image size (e.g. 768x768) also on gldv2x for a few epochs. I used Adam optimizer and cosine annealing schedule in every step. I used same augmentation as *All Data Are Ext*, which proved to be very efficient:\n\n```\ncfg.train_aug = A.Compose([\n        A.HorizontalFlip(p=0.5),\n        A.ImageCompression(quality_lower=99, quality_upper=100),\n        A.ShiftScaleRotate(shift_limit=0.2, scale_limit=0.2, rotate_limit=10, border_mode=0, p=0.7),\n        A.Resize(image_size, image_size),\n        A.Cutout(max_h_size=int(image_size * 0.4), max_w_size=int(image_size * 0.4), num_holes=1, p=0.5),\n    ])\n\ncfg.val_aug = A.Compose([A.Resize(image_size, image_size),])\n\n```\n\n## Model Architectures\n\nFor me this is the most fun part in every competition, to design and explore model architectures. I like the idea of models solving multiple parts of the problem at hand in end2end fashion. In landmark recognition it is crucial to leverage both, global and local image information. While historically this was solved in a two-stage fashion by training a global descriptor for preliminary recognition on the one hand and extracting local descriptors for post reranking using spatial verification on the other hand. Recently attempts have been made to train local and global descriptors simultaneously using a single model (DELF/ DELG). But that still uses a 2nd stage spatial verification. \nI implemented two architectures that combine local and global descriptor into a single image embedding. Both use an effientnet encoder and an sub-arcface head with dynamic margins which was shown by team *All Data Are Ext* in 2020 to outperform classic arcface.\n\n### DOLG\n\nThe very recently released DOLG paper, goes on step further claiming to train local and global descriptors also simultaneously but additionally fusing them within the model into a single descriptor. Unfortunately/ fortunately it was so new that no code was released so far and I needed to implement deep orthogonal fusion by myself. Luckily the paper gives enough details and I was able to implement it quickly. The following shows the final architecture:\n\n![DOLG-EffiencentNet](https://i.imgur.com/9FiSEet.png)\n\n### Hybrid Swin Transformer\n\nAnother architecture I explored was a Hybrid-Transformer, half a CNN encoder with a transformer ending. The idea behind is that transformers are like graph neural nets on image patches connecting distant local features with each other. A property CNNs are lacking, but which are very important for landmark recognition/ retrieval. However, pretrained image transformers are mostly done on 224 or 384 image size, and I wanted to go bigger. Also the encoding of each patch is of a higher quality when using a sophisticated CNN for the patch embedding. Luckily the awesome timm repository gives you the tool set to glue together a CNN and an image transformer and I ended with an architecture like shown below. In particular I used Swin transformer [3], as I think the moving window approach captures more diverse local descriptors and EfficientNet for encoding.\n\n![Hybrid-Swin-Transformer](https://i.imgur.com/2iXNuBA.png)\n\nHowever, for training several problems arise: Firstly the pretrained weights of image transformer and CNN are trained separately, hence not fitting very well together, resulting in NaNs while training. Secondly, since image transformer require a fixed input size it's a bit more difficult to follow a training routine with increasing image size in each step. I found that the following procedure fixes the issues and gives a good final model, and especially step 3 is important to make the Hybrid-Swin-Transformer working\n\n1. train only image transformer on 224x224 size\n2. Exchange original patch embedding module with block 0,1,2 from Effnet encoder, \n3. freeze image transformer and sub-arcface head and train for 1 epoch on 448x448, to let the effnet encoder “adjust”\n4. unfreeze image transformer and head and train for 30-40 epochs\n5. add block 3 from effnet encoder and finetune a few epochs on 896x896\n\nFor two of my models I used the “stride trick”, i.e. putting a stride of (1,1) in the first conv layer to effectively double the used resolution instead of double the actual image size, which I think is better for small original images.\nI also retrained a few models from 2nd and 3rd place solutions from 2020 recognition competition following their provided git repositories and instructions, and evaluated possible ensembles on my local cross validation. Models form All Data Ext team where quite low correlated to mine and added a nice diversification benefit\n\nIn my final ensemble I had the following models:\n- DOLG-Efficentnet b5 (768x768)\n- DOLG-Efficentnet b6 (768x768)\n- DOLG-Efficentnet b7 (448x448) stride 1\n- Hybrid SwinBase224-Efficentnet b5 (448x448) stride 1\n- Hybrid SwinBase224-Efficentnet b3 (896x896)\n- Hybrid SwinBase384-Efficentnet b6 (384,384) stride 1\n- All Data Ext 2020 Efficientnet b6 (512x512) \n- All Data Ext 2020 EfficientNet b3 (768x768)\n\nPost-processing\nI literally used the same approach (developed by @philippsinger) as in our last year solution. you can find it here:\n\nhttps://www.kaggle.com/c/landmark-recognition-2020/discussion/187821\n\nFor more details please refer to the paper on arxiv. Thank you for reading.\n\n[1] https://arxiv.org/abs/2108.02927 (DOLG)\n[2] https://github.com/cvdfoundation/google-landmark\n[3] https://arxiv.org/abs/2103.14030 (Swin)\n\nCode: https://github.com/ChristofHenkel/kaggle-landmark-2021-1st-place (in progress) \nPaper: submitted to arxiv, waiting for acception. In the meanwhile I uploaded to https://github.com/ChristofHenkel/kaggle-landmark-2021-1st-place/blob/main/GLR_2021_1st_place.pdf\n\n",
      "votes": 184
    },
    {
      "id": 1682949,
      "postDate": "2022-02-09T13:36:57.833Z",
      "content": "<p>congrats <a href=\"https://www.kaggle.com/christofhenkel\" target=\"_blank\">@christofhenkel</a> 🤩</p>",
      "rawMarkdown": "congrats @christofhenkel 🤩",
      "votes": 1
    },
    {
      "id": 1539952,
      "postDate": "2021-10-10T02:40:36.580Z",
      "content": "<p>I'm glad to know your solution. it gave me great inspiration.👍👍<br>\nThank you😃</p>",
      "rawMarkdown": "I'm glad to know your solution. it gave me great inspiration.👍👍\nThank you😃",
      "votes": 1
    },
    {
      "id": 1539070,
      "postDate": "2021-10-09T05:14:29.760Z",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/christofhenkel\" target=\"_blank\">@christofhenkel</a> on winning the comp and also becoming 1st ranked GM.<br>\nTraining these big models with so much big data is not that easy, so what was your experiment setup or your training recipe. I mean did you keep experimenting with different architecture, hyperparameters manually or were you using some tool like optuna or something to select the right hyperparameters and all.</p>",
      "rawMarkdown": "Congratulations @christofhenkel on winning the comp and also becoming 1st ranked GM.\nTraining these big models with so much big data is not that easy, so what was your experiment setup or your training recipe. I mean did you keep experimenting with different architecture, hyperparameters manually or were you using some tool like optuna or something to select the right hyperparameters and all.",
      "votes": 1,
      "replies": [
        {
          "id": 1539301,
          "postDate": "2021-10-09T09:46:21.493Z",
          "content": "<p>To be honest, given the tighter timeline this year, there was no time for hyperparameter tuning. I mainly used hyperparameters from last year/ experience and focused on architectural changes. However, and I think thats important, I spend some time to implement local validation of retrieval score at the end of every epoch, so that I can asses model quality.</p>",
          "rawMarkdown": "To be honest, given the tighter timeline this year, there was no time for hyperparameter tuning. I mainly used hyperparameters from last year/ experience and focused on architectural changes. However, and I think thats important, I spend some time to implement local validation of retrieval score at the end of every epoch, so that I can asses model quality.",
          "votes": 4
        }
      ]
    },
    {
      "id": 1542606,
      "postDate": "2021-10-12T17:17:38.240Z",
      "content": "<p>Awesome models! Congrats on the double win. Thanks for sharing your models and solution summary.</p>",
      "rawMarkdown": "Awesome models! Congrats on the double win. Thanks for sharing your models and solution summary.",
      "votes": 2
    },
    {
      "id": 1540145,
      "postDate": "2021-10-10T07:23:25.320Z",
      "content": "<p>Much congratulations <a href=\"https://www.kaggle.com/christofhenkel\" target=\"_blank\">@christofhenkel</a> for this wonderful achievement!</p>\n<p>In your paper, you mentioned that you trained your models on 8xV100 NVIDIA GPUs with distributed data parallel; did you use a cloud provider, or did you have all the GPUs in hand? For new Kagglers, what GPU(s) would you recommend us getting?</p>\n<p>Also, how long did it take to train your entire ensemble, with distributed data parallel?</p>",
      "rawMarkdown": "Much congratulations @christofhenkel for this wonderful achievement!\n\nIn your paper, you mentioned that you trained your models on 8xV100 NVIDIA GPUs with distributed data parallel; did you use a cloud provider, or did you have all the GPUs in hand? For new Kagglers, what GPU(s) would you recommend us getting?\n\nAlso, how long did it take to train your entire ensemble, with distributed data parallel?",
      "votes": 2,
      "replies": [
        {
          "id": 1540998,
          "postDate": "2021-10-11T06:02:46.180Z",
          "content": "<p>Luckily, I have access to a DGX-1 (8xV100 32GB) at work. When I started kaggling I first used colab and then relatively quickly build my own Desktop-PC with 2-GPUs (GTX 1080Ti at that time) following some youtube/ medium suggestions for parts. </p>",
          "rawMarkdown": "Luckily, I have access to a DGX-1 (8xV100 32GB) at work. When I started kaggling I first used colab and then relatively quickly build my own Desktop-PC with 2-GPUs (GTX 1080Ti at that time) following some youtube/ medium suggestions for parts. ",
          "votes": 5
        }
      ]
    },
    {
      "id": 1672170,
      "postDate": "2022-02-01T22:15:01.950Z",
      "content": "<p>How did u label previous test sets for your validation set?</p>",
      "rawMarkdown": "How did u label previous test sets for your validation set?",
      "replies": [
        {
          "id": 1682983,
          "postDate": "2022-02-09T13:53:57.200Z",
          "content": "<p>labels of previous test sets were released after previous competitions</p>",
          "rawMarkdown": "labels of previous test sets were released after previous competitions"
        }
      ]
    },
    {
      "id": 1656359,
      "postDate": "2022-01-19T09:42:45.713Z",
      "content": "<p>What loss function have u used?</p>",
      "rawMarkdown": "What loss function have u used?\n",
      "replies": [
        {
          "id": 1656884,
          "postDate": "2022-01-19T18:21:33.687Z",
          "content": "<p>sub-center arcface</p>\n<p><a href=\"https://ibug.doc.ic.ac.uk/media/uploads/documents/eccv_1445.pdf\" target=\"_blank\">https://ibug.doc.ic.ac.uk/media/uploads/documents/eccv_1445.pdf</a></p>",
          "rawMarkdown": "sub-center arcface\n\nhttps://ibug.doc.ic.ac.uk/media/uploads/documents/eccv_1445.pdf"
        },
        {
          "id": 1661324,
          "postDate": "2022-01-23T11:40:18.133Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 1585935,
      "postDate": "2021-11-17T16:56:33.603Z",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/christofhenkel\" target=\"_blank\">@christofhenkel</a>, thank you for sharing and your models are really fascinating, especially the hybrid model, and I am trying to replicate it. However, there are some points that are a bit unclear to me, I hope you could take a bit of time to explain them. Let's focus on the \"Hybrid SwinBase224-Efficentnet b3 (896x896)\" model, I think the others are similar.</p>\n<ol>\n<li><p>In your write-up, you said \"Exchange original patch embedding module with block 0,1,2 from Effnet encoder\",  I think it should be \"blocks 0, 1\" only because the output of it has the size of (BS, 32, 224, 224), isn't it? Accordingly, in step 5, \"add block 3 from effnet encoder and finetune a few epochs on 896x896\", block 2 should be added, not 3, if blocks 0, 1 are used.</p></li>\n<li><p>As I experimented, the output size of the first 2 blocks of EfficientNet-b5 is (BS, 32, 224, 224), but the input size of the Swin Transformer should be (BS, 3136, 128), this is the output size after the \"patch_embed\" layer. </p></li>\n</ol>\n<pre><code>(patch_embed): PatchEmbed(\n    (proj): Conv2d(32, 128, kernel_size=(4, 4), stride=(4, 4))\n    (norm): LayerNorm((128,), eps=1e-05, elementwise_affine=True)\n  )\n</code></pre>\n<p>You mentioned that the output of the 2-block of EfficientNet-b5 is flattened, added with positional embeddings, and projected. As far as I understand, \"flattened\" here is with the patch (4 x 4) level (not pixel level (1 x 1)), so we can't use nn.Flatten(), am I correct? After flattening and transposing with patches, the input of the Swin transformer should be (BS, 3136, 32), and now you projected it into the (BS, 3136, 128) by nn.Linear(32, 128) to fit the input dim of the transformer. If all of this is correctly understood from my side, one question is, why don't you simply input the output of the 2-block EfficientNet-B5 into the Swin transformer at the beginning (before the \"patch_embed\" layer)?</p>\n<ol>\n<li><p>You mentioned that you use ArcFace head with a \"dynamic\" margin, can I ask what \"dynamic\" here mean? Is it changed over training steps, epochs, or something else?</p></li>\n<li><p>The last question is the \"virtual_token\", I am curious how to add and initialize it into the embeddings, is it simply like adding one 128-dim vector into the (BS, 3136, 128) to become (BS, 3137, 128)?</p></li>\n</ol>\n<p>My questions have been super long and I hope it doesn't bother you.</p>\n<p>Thank you!</p>",
      "rawMarkdown": "Hi @christofhenkel, thank you for sharing and your models are really fascinating, especially the hybrid model, and I am trying to replicate it. However, there are some points that are a bit unclear to me, I hope you could take a bit of time to explain them. Let's focus on the \"Hybrid SwinBase224-Efficentnet b3 (896x896)\" model, I think the others are similar.\n\n1. In your write-up, you said \"Exchange original patch embedding module with block 0,1,2 from Effnet encoder\",  I think it should be \"blocks 0, 1\" only because the output of it has the size of (BS, 32, 224, 224), isn't it? Accordingly, in step 5, \"add block 3 from effnet encoder and finetune a few epochs on 896x896\", block 2 should be added, not 3, if blocks 0, 1 are used.\n\n2. As I experimented, the output size of the first 2 blocks of EfficientNet-b5 is (BS, 32, 224, 224), but the input size of the Swin Transformer should be (BS, 3136, 128), this is the output size after the \"patch_embed\" layer. \n```\n(patch_embed): PatchEmbed(\n    (proj): Conv2d(32, 128, kernel_size=(4, 4), stride=(4, 4))\n    (norm): LayerNorm((128,), eps=1e-05, elementwise_affine=True)\n  )\n```\nYou mentioned that the output of the 2-block of EfficientNet-b5 is flattened, added with positional embeddings, and projected. As far as I understand, \"flattened\" here is with the patch (4 x 4) level (not pixel level (1 x 1)), so we can't use nn.Flatten(), am I correct? After flattening and transposing with patches, the input of the Swin transformer should be (BS, 3136, 32), and now you projected it into the (BS, 3136, 128) by nn.Linear(32, 128) to fit the input dim of the transformer. If all of this is correctly understood from my side, one question is, why don't you simply input the output of the 2-block EfficientNet-B5 into the Swin transformer at the beginning (before the \"patch_embed\" layer)?\n\n3. You mentioned that you use ArcFace head with a \"dynamic\" margin, can I ask what \"dynamic\" here mean? Is it changed over training steps, epochs, or something else?\n\n4. The last question is the \"virtual_token\", I am curious how to add and initialize it into the embeddings, is it simply like adding one 128-dim vector into the (BS, 3136, 128) to become (BS, 3137, 128)?\n\nMy questions have been super long and I hope it doesn't bother you.\n\nThank you!"
    },
    {
      "id": 1572295,
      "postDate": "2021-11-05T15:28:54.843Z",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/christofhenkel\" target=\"_blank\">@christofhenkel</a> for the double win, and also thanks for sharing your ideas.<br>\nI have some questions about the DOLF Efficientnet. In the original paper of DOLG, the Local Branch consists of Multi-Atrous layers and a pooling layer and is followed by a self-attention module after concatenation. But in your design, do you drop the pooling layer and resample the size to half of its original feature map's size?</p>",
      "rawMarkdown": "Congrats @christofhenkel for the double win, and also thanks for sharing your ideas.\nI have some questions about the DOLF Efficientnet. In the original paper of DOLG, the Local Branch consists of Multi-Atrous layers and a pooling layer and is followed by a self-attention module after concatenation. But in your design, do you drop the pooling layer and resample the size to half of its original feature map's size?",
      "replies": [
        {
          "id": 1572475,
          "postDate": "2021-11-05T17:42:18.110Z",
          "content": "<p>yes I skipped the pooling layer, as it was not transparent to me, how/ what they pool. What do you mean with \"resample the size to half of its original feature map's size\"? </p>\n<pre><code>class MultiAtrousModule(nn.Module):\n    def __init__(self, in_chans, out_chans, dilations):\n        super(MultiAtrousModule, self).__init__()\n\n        self.d6 = nn.Conv2d(in_chans, 512, kernel_size=3, dilation=dilations[0],padding='same')\n        self.d12 = nn.Conv2d(in_chans, 512, kernel_size=3, dilation=dilations[1],padding='same')\n        self.d18 = nn.Conv2d(in_chans, 512, kernel_size=3, dilation=dilations[2],padding='same')\n        self.conv1 = nn.Conv2d(512 * 3, out_chans, kernel_size=1)\n        self.relu = nn.ReLU()\n\n    def forward(self,x):\n\n        x6 = self.d6(x)\n        x12 = self.d12(x)\n        x18 = self.d18(x)\n        x = torch.cat([x6,x12,x18],dim=1)\n        x = self.conv1(x)\n        x = self.relu(x)\n        return x\n</code></pre>\n<p>taking input tensor of (bs,w,h,in_chans) this module outputs(bs,w,h,out_chans)</p>",
          "rawMarkdown": "yes I skipped the pooling layer, as it was not transparent to me, how/ what they pool. What do you mean with \"resample the size to half of its original feature map's size\"? \n\n```\nclass MultiAtrousModule(nn.Module):\n    def __init__(self, in_chans, out_chans, dilations):\n        super(MultiAtrousModule, self).__init__()\n        \n        self.d6 = nn.Conv2d(in_chans, 512, kernel_size=3, dilation=dilations[0],padding='same')\n        self.d12 = nn.Conv2d(in_chans, 512, kernel_size=3, dilation=dilations[1],padding='same')\n        self.d18 = nn.Conv2d(in_chans, 512, kernel_size=3, dilation=dilations[2],padding='same')\n        self.conv1 = nn.Conv2d(512 * 3, out_chans, kernel_size=1)\n        self.relu = nn.ReLU()\n        \n    def forward(self,x):\n        \n        x6 = self.d6(x)\n        x12 = self.d12(x)\n        x18 = self.d18(x)\n        x = torch.cat([x6,x12,x18],dim=1)\n        x = self.conv1(x)\n        x = self.relu(x)\n        return x\n```\n\ntaking input tensor of (bs,w,h,in_chans) this module outputs(bs,w,h,out_chans)\n",
          "votes": 1
        },
        {
          "id": 1572764,
          "postDate": "2021-11-05T23:22:07.990Z",
          "content": "<p>Thanks for your fast reply. The \"resample\" thing is because in the \"Figure 1“  model architecture of DOLG Efficientnet b5 in your paper (<a href=\"https://github.com/ChristofHenkel/kaggle-landmark-2021-1st-place/blob/main/GLR_2021_1st_place.pdf\" target=\"_blank\">paper</a>), the local descriptor is (24x24x1024) but not  (48x48x1024). Is it a typo? or a reduction block is added at the projection and spatial attention?</p>",
          "rawMarkdown": "Thanks for your fast reply. The \"resample\" thing is because in the \"Figure 1“  model architecture of DOLG Efficientnet b5 in your paper ([paper] (https://github.com/ChristofHenkel/kaggle-landmark-2021-1st-place/blob/main/GLR_2021_1st_place.pdf)), the local descriptor is (24x24x1024) but not  (48x48x1024). Is it a typo? or a reduction block is added at the projection and spatial attention?\n"
        },
        {
          "id": 1572908,
          "postDate": "2021-11-06T04:49:01.693Z",
          "content": "<p>its a typo, thanks for noting</p>",
          "rawMarkdown": "its a typo, thanks for noting"
        }
      ]
    },
    {
      "id": 1558204,
      "postDate": "2021-10-26T06:10:01.180Z",
      "content": "<p>Congrats! I learned a lot from your solution.</p>",
      "rawMarkdown": "Congrats! I learned a lot from your solution."
    },
    {
      "id": 1557241,
      "postDate": "2021-10-25T14:19:38.030Z",
      "content": "<p><a href=\"https://www.kaggle.com/christofhenkel\" target=\"_blank\">@christofhenkel</a> Congratulations on winning both landmark competitions and becoming #1 in leaderboard!</p>\n<p>If you don't mind me asking, when is the code for your solution going to be available (it's currently in progress)? I understand that you're probably busy doing clean-up, but I'm just wondering when you think it will be ready to be uploaded. Thank you!</p>",
      "rawMarkdown": "@christofhenkel Congratulations on winning both landmark competitions and becoming #1 in leaderboard!\n\nIf you don't mind me asking, when is the code for your solution going to be available (it's currently in progress)? I understand that you're probably busy doing clean-up, but I'm just wondering when you think it will be ready to be uploaded. Thank you!"
    },
    {
      "id": 1551516,
      "postDate": "2021-10-20T17:27:14.543Z",
      "content": "<p>Wonderful!</p>",
      "rawMarkdown": "Wonderful!"
    },
    {
      "id": 1551069,
      "postDate": "2021-10-20T09:02:44.883Z",
      "content": "<p>Excellent!</p>",
      "rawMarkdown": "Excellent!"
    },
    {
      "id": 1550799,
      "postDate": "2021-10-20T03:07:11.637Z",
      "content": "<p>It's amazing!</p>",
      "rawMarkdown": "It's amazing!"
    },
    {
      "id": 1549571,
      "postDate": "2021-10-19T04:18:01.410Z",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/christofhenkel\" target=\"_blank\">@christofhenkel</a> !!!.👍👍👍👍👍👍👍<br>\nI have been wondering how you choose LR schedulers and optimizers in such a large data set (assuming you have enough time ), and how you can quickly verify your ideas, and whether you would consider using a small number of data sets for validation first ? <br>\nOr could you recommend some papers so that I can find the answer by myself.  </p>",
      "rawMarkdown": "Congratulations @christofhenkel !!!.👍👍👍👍👍👍👍\nI have been wondering how you choose LR schedulers and optimizers in such a large data set (assuming you have enough time ), and how you can quickly verify your ideas, and whether you would consider using a small number of data sets for validation first ? \nOr could you recommend some papers so that I can find the answer by myself.  ",
      "replies": [
        {
          "id": 1549638,
          "postDate": "2021-10-19T05:36:58.470Z",
          "content": "<p>I mainly used hyperparameters from last year and experience from other competitions. But in general its a good idea to use small image size to roughly verify different optimizers and schedule. Also, as you say, only using a portion of the dataset will further improve time-efficiency for those choices.</p>",
          "rawMarkdown": " I mainly used hyperparameters from last year and experience from other competitions. But in general its a good idea to use small image size to roughly verify different optimizers and schedule. Also, as you say, only using a portion of the dataset will further improve time-efficiency for those choices."
        }
      ]
    },
    {
      "id": 1543113,
      "postDate": "2021-10-13T06:42:22.907Z",
      "content": "<p>Nice work!</p>",
      "rawMarkdown": "Nice work!"
    },
    {
      "id": 1543068,
      "postDate": "2021-10-13T05:46:31.773Z",
      "content": "<p>Congratulation on the win. Also this is an work</p>",
      "rawMarkdown": "Congratulation on the win. Also this is an work"
    },
    {
      "id": 1542269,
      "postDate": "2021-10-12T10:12:11.173Z",
      "content": "<p>Amazing work. Congratulations <a href=\"https://www.kaggle.com/christofhenkel\" target=\"_blank\">@christofhenkel</a>.</p>",
      "rawMarkdown": "Amazing work. Congratulations @christofhenkel."
    },
    {
      "id": 1540422,
      "postDate": "2021-10-10T12:59:23.273Z",
      "content": "<p><a href=\"https://www.kaggle.com/christofhenkel\" target=\"_blank\">@christofhenkel</a> Congrats on your win and also for ranking on on the Competition Pillar!  Thank you for your write up, I found it very useful how you approach a problem with a tight timeline - leveraging previous work (and thanks to for those links).  </p>",
      "rawMarkdown": "@christofhenkel Congrats on your win and also for ranking on on the Competition Pillar!  Thank you for your write up, I found it very useful how you approach a problem with a tight timeline - leveraging previous work (and thanks to for those links).  "
    },
    {
      "id": 1539996,
      "postDate": "2021-10-10T04:31:05.650Z",
      "content": "<p>Congratulations!. Thanks a lot for sharing.</p>",
      "rawMarkdown": "Congratulations!. Thanks a lot for sharing."
    },
    {
      "id": 1538043,
      "postDate": "2021-10-08T02:55:45.443Z",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/christofhenkel\" target=\"_blank\">@christofhenkel</a>!<br>\nI also tried progressive training but the loss increased for the first few epochs. Wanted to know if this is expected behaviour or something wrong with the implementation?</p>",
      "rawMarkdown": "Congratulations @christofhenkel!\nI also tried progressive training but the loss increased for the first few epochs. Wanted to know if this is expected behaviour or something wrong with the implementation?",
      "replies": [
        {
          "id": 1538099,
          "postDate": "2021-10-08T04:26:18.627Z",
          "content": "<p>It depends on several factors, especially if you change more than just image size. Also when going from GLDv2c to GLDV2x, its expected to have a higher loss, as data becomes more noisy. If you only change img size, loss should converge to old level pretty fast.</p>",
          "rawMarkdown": "It depends on several factors, especially if you change more than just image size. Also when going from GLDv2c to GLDV2x, its expected to have a higher loss, as data becomes more noisy. If you only change img size, loss should converge to old level pretty fast.",
          "votes": 3
        }
      ]
    },
    {
      "id": 3275451,
      "postDate": "2025-08-26T14:54:20.903Z",
      "content": "<p>Congratulations</p>",
      "rawMarkdown": "Congratulations"
    },
    {
      "id": 1561494,
      "postDate": "2021-10-27T15:26:00.907Z",
      "content": "<p>Congratulations and thanks for sharing</p>",
      "rawMarkdown": "Congratulations and thanks for sharing"
    },
    {
      "id": 1542994,
      "postDate": "2021-10-13T03:55:56.773Z",
      "content": "<p>It's amazing! Thank you for sharing.</p>",
      "rawMarkdown": "It's amazing! Thank you for sharing."
    }
  ],
  "comments": [
    {
      "id": 1682949,
      "author_name": "Aruna S",
      "author_url": "",
      "post_date": "2022-02-09T13:36:57.833000",
      "content": "<p>congrats <a href=\"https://www.kaggle.com/christofhenkel\" target=\"_blank\">@christofhenkel</a> 🤩</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1539952,
      "author_name": "J.H.Lee",
      "author_url": "",
      "post_date": "2021-10-10T02:40:36.580000",
      "content": "<p>I'm glad to know your solution. it gave me great inspiration.👍👍<br>\nThank you😃</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1539070,
      "author_name": "DeepUnderstanding",
      "author_url": "",
      "post_date": "2021-10-09T05:14:29.760000",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/christofhenkel\" target=\"_blank\">@christofhenkel</a> on winning the comp and also becoming 1st ranked GM.<br>\nTraining these big models with so much big data is not that easy, so what was your experiment setup or your training recipe. I mean did you keep experimenting with different architecture, hyperparameters manually or were you using some tool like optuna or something to select the right hyperparameters and all.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1539301,
          "author_name": "Dieter",
          "author_url": "",
          "post_date": "2021-10-09T09:46:21.493000",
          "content": "<p>To be honest, given the tighter timeline this year, there was no time for hyperparameter tuning. I mainly used hyperparameters from last year/ experience and focused on architectural changes. However, and I think thats important, I spend some time to implement local validation of retrieval score at the end of every epoch, so that I can asses model quality.</p>",
          "votes": 4,
          "replies": []
        }
      ]
    },
    {
      "id": 1542606,
      "author_name": "Chris Deotte",
      "author_url": "",
      "post_date": "2021-10-12T17:17:38.240000",
      "content": "<p>Awesome models! Congrats on the double win. Thanks for sharing your models and solution summary.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1540145,
      "author_name": "Coderrexe",
      "author_url": "",
      "post_date": "2021-10-10T07:23:25.320000",
      "content": "<p>Much congratulations <a href=\"https://www.kaggle.com/christofhenkel\" target=\"_blank\">@christofhenkel</a> for this wonderful achievement!</p>\n<p>In your paper, you mentioned that you trained your models on 8xV100 NVIDIA GPUs with distributed data parallel; did you use a cloud provider, or did you have all the GPUs in hand? For new Kagglers, what GPU(s) would you recommend us getting?</p>\n<p>Also, how long did it take to train your entire ensemble, with distributed data parallel?</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1540998,
          "author_name": "Dieter",
          "author_url": "",
          "post_date": "2021-10-11T06:02:46.180000",
          "content": "<p>Luckily, I have access to a DGX-1 (8xV100 32GB) at work. When I started kaggling I first used colab and then relatively quickly build my own Desktop-PC with 2-GPUs (GTX 1080Ti at that time) following some youtube/ medium suggestions for parts. </p>",
          "votes": 5,
          "replies": []
        }
      ]
    },
    {
      "id": 1672170,
      "author_name": "zelenadinja",
      "author_url": "",
      "post_date": "2022-02-01T22:15:01.950000",
      "content": "<p>How did u label previous test sets for your validation set?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1682983,
          "author_name": "Dieter",
          "author_url": "",
          "post_date": "2022-02-09T13:53:57.200000",
          "content": "<p>labels of previous test sets were released after previous competitions</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1656359,
      "author_name": "zelenadinja",
      "author_url": "",
      "post_date": "2022-01-19T09:42:45.713000",
      "content": "<p>What loss function have u used?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1656884,
          "author_name": "Dieter",
          "author_url": "",
          "post_date": "2022-01-19T18:21:33.687000",
          "content": "<p>sub-center arcface</p>\n<p><a href=\"https://ibug.doc.ic.ac.uk/media/uploads/documents/eccv_1445.pdf\" target=\"_blank\">https://ibug.doc.ic.ac.uk/media/uploads/documents/eccv_1445.pdf</a></p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1661324,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-01-23T11:40:18.133000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1585935,
      "author_name": "Minh Tri Phan",
      "author_url": "",
      "post_date": "2021-11-17T16:56:33.603000",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/christofhenkel\" target=\"_blank\">@christofhenkel</a>, thank you for sharing and your models are really fascinating, especially the hybrid model, and I am trying to replicate it. However, there are some points that are a bit unclear to me, I hope you could take a bit of time to explain them. Let's focus on the \"Hybrid SwinBase224-Efficentnet b3 (896x896)\" model, I think the others are similar.</p>\n<ol>\n<li><p>In your write-up, you said \"Exchange original patch embedding module with block 0,1,2 from Effnet encoder\",  I think it should be \"blocks 0, 1\" only because the output of it has the size of (BS, 32, 224, 224), isn't it? Accordingly, in step 5, \"add block 3 from effnet encoder and finetune a few epochs on 896x896\", block 2 should be added, not 3, if blocks 0, 1 are used.</p></li>\n<li><p>As I experimented, the output size of the first 2 blocks of EfficientNet-b5 is (BS, 32, 224, 224), but the input size of the Swin Transformer should be (BS, 3136, 128), this is the output size after the \"patch_embed\" layer. </p></li>\n</ol>\n<pre><code>(patch_embed): PatchEmbed(\n    (proj): Conv2d(32, 128, kernel_size=(4, 4), stride=(4, 4))\n    (norm): LayerNorm((128,), eps=1e-05, elementwise_affine=True)\n  )\n</code></pre>\n<p>You mentioned that the output of the 2-block of EfficientNet-b5 is flattened, added with positional embeddings, and projected. As far as I understand, \"flattened\" here is with the patch (4 x 4) level (not pixel level (1 x 1)), so we can't use nn.Flatten(), am I correct? After flattening and transposing with patches, the input of the Swin transformer should be (BS, 3136, 32), and now you projected it into the (BS, 3136, 128) by nn.Linear(32, 128) to fit the input dim of the transformer. If all of this is correctly understood from my side, one question is, why don't you simply input the output of the 2-block EfficientNet-B5 into the Swin transformer at the beginning (before the \"patch_embed\" layer)?</p>\n<ol>\n<li><p>You mentioned that you use ArcFace head with a \"dynamic\" margin, can I ask what \"dynamic\" here mean? Is it changed over training steps, epochs, or something else?</p></li>\n<li><p>The last question is the \"virtual_token\", I am curious how to add and initialize it into the embeddings, is it simply like adding one 128-dim vector into the (BS, 3136, 128) to become (BS, 3137, 128)?</p></li>\n</ol>\n<p>My questions have been super long and I hope it doesn't bother you.</p>\n<p>Thank you!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1572295,
      "author_name": "haha6ybisme",
      "author_url": "",
      "post_date": "2021-11-05T15:28:54.843000",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/christofhenkel\" target=\"_blank\">@christofhenkel</a> for the double win, and also thanks for sharing your ideas.<br>\nI have some questions about the DOLF Efficientnet. In the original paper of DOLG, the Local Branch consists of Multi-Atrous layers and a pooling layer and is followed by a self-attention module after concatenation. But in your design, do you drop the pooling layer and resample the size to half of its original feature map's size?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1572475,
          "author_name": "Dieter",
          "author_url": "",
          "post_date": "2021-11-05T17:42:18.110000",
          "content": "<p>yes I skipped the pooling layer, as it was not transparent to me, how/ what they pool. What do you mean with \"resample the size to half of its original feature map's size\"? </p>\n<pre><code>class MultiAtrousModule(nn.Module):\n    def __init__(self, in_chans, out_chans, dilations):\n        super(MultiAtrousModule, self).__init__()\n\n        self.d6 = nn.Conv2d(in_chans, 512, kernel_size=3, dilation=dilations[0],padding='same')\n        self.d12 = nn.Conv2d(in_chans, 512, kernel_size=3, dilation=dilations[1],padding='same')\n        self.d18 = nn.Conv2d(in_chans, 512, kernel_size=3, dilation=dilations[2],padding='same')\n        self.conv1 = nn.Conv2d(512 * 3, out_chans, kernel_size=1)\n        self.relu = nn.ReLU()\n\n    def forward(self,x):\n\n        x6 = self.d6(x)\n        x12 = self.d12(x)\n        x18 = self.d18(x)\n        x = torch.cat([x6,x12,x18],dim=1)\n        x = self.conv1(x)\n        x = self.relu(x)\n        return x\n</code></pre>\n<p>taking input tensor of (bs,w,h,in_chans) this module outputs(bs,w,h,out_chans)</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1572764,
          "author_name": "haha6ybisme",
          "author_url": "",
          "post_date": "2021-11-05T23:22:07.990000",
          "content": "<p>Thanks for your fast reply. The \"resample\" thing is because in the \"Figure 1“  model architecture of DOLG Efficientnet b5 in your paper (<a href=\"https://github.com/ChristofHenkel/kaggle-landmark-2021-1st-place/blob/main/GLR_2021_1st_place.pdf\" target=\"_blank\">paper</a>), the local descriptor is (24x24x1024) but not  (48x48x1024). Is it a typo? or a reduction block is added at the projection and spatial attention?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1572908,
          "author_name": "Dieter",
          "author_url": "",
          "post_date": "2021-11-06T04:49:01.693000",
          "content": "<p>its a typo, thanks for noting</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1558204,
      "author_name": "Regulus Bai",
      "author_url": "",
      "post_date": "2021-10-26T06:10:01.180000",
      "content": "<p>Congrats! I learned a lot from your solution.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1557241,
      "author_name": "JGeoff",
      "author_url": "",
      "post_date": "2021-10-25T14:19:38.030000",
      "content": "<p><a href=\"https://www.kaggle.com/christofhenkel\" target=\"_blank\">@christofhenkel</a> Congratulations on winning both landmark competitions and becoming #1 in leaderboard!</p>\n<p>If you don't mind me asking, when is the code for your solution going to be available (it's currently in progress)? I understand that you're probably busy doing clean-up, but I'm just wondering when you think it will be ready to be uploaded. Thank you!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1551516,
      "author_name": "Aashish Bhandari",
      "author_url": "",
      "post_date": "2021-10-20T17:27:14.543000",
      "content": "<p>Wonderful!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1551069,
      "author_name": "Nidhin Thomas",
      "author_url": "",
      "post_date": "2021-10-20T09:02:44.883000",
      "content": "<p>Excellent!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1550799,
      "author_name": "Andy Chen",
      "author_url": "",
      "post_date": "2021-10-20T03:07:11.637000",
      "content": "<p>It's amazing!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1549571,
      "author_name": "lafe",
      "author_url": "",
      "post_date": "2021-10-19T04:18:01.410000",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/christofhenkel\" target=\"_blank\">@christofhenkel</a> !!!.👍👍👍👍👍👍👍<br>\nI have been wondering how you choose LR schedulers and optimizers in such a large data set (assuming you have enough time ), and how you can quickly verify your ideas, and whether you would consider using a small number of data sets for validation first ? <br>\nOr could you recommend some papers so that I can find the answer by myself.  </p>",
      "votes": 0,
      "replies": [
        {
          "id": 1549638,
          "author_name": "Dieter",
          "author_url": "",
          "post_date": "2021-10-19T05:36:58.470000",
          "content": "<p>I mainly used hyperparameters from last year and experience from other competitions. But in general its a good idea to use small image size to roughly verify different optimizers and schedule. Also, as you say, only using a portion of the dataset will further improve time-efficiency for those choices.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1543113,
      "author_name": "Faizal Karim",
      "author_url": "",
      "post_date": "2021-10-13T06:42:22.907000",
      "content": "<p>Nice work!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1543068,
      "author_name": "Shlok Ramteke",
      "author_url": "",
      "post_date": "2021-10-13T05:46:31.773000",
      "content": "<p>Congratulation on the win. Also this is an work</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1542269,
      "author_name": "Mukharbek Organokov",
      "author_url": "",
      "post_date": "2021-10-12T10:12:11.173000",
      "content": "<p>Amazing work. Congratulations <a href=\"https://www.kaggle.com/christofhenkel\" target=\"_blank\">@christofhenkel</a>.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1540422,
      "author_name": "YvonneF",
      "author_url": "",
      "post_date": "2021-10-10T12:59:23.273000",
      "content": "<p><a href=\"https://www.kaggle.com/christofhenkel\" target=\"_blank\">@christofhenkel</a> Congrats on your win and also for ranking on on the Competition Pillar!  Thank you for your write up, I found it very useful how you approach a problem with a tight timeline - leveraging previous work (and thanks to for those links).  </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1539996,
      "author_name": "Saleesh",
      "author_url": "",
      "post_date": "2021-10-10T04:31:05.650000",
      "content": "<p>Congratulations!. Thanks a lot for sharing.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1538043,
      "author_name": "Debarshi Chanda",
      "author_url": "",
      "post_date": "2021-10-08T02:55:45.443000",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/christofhenkel\" target=\"_blank\">@christofhenkel</a>!<br>\nI also tried progressive training but the loss increased for the first few epochs. Wanted to know if this is expected behaviour or something wrong with the implementation?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1538099,
          "author_name": "Dieter",
          "author_url": "",
          "post_date": "2021-10-08T04:26:18.627000",
          "content": "<p>It depends on several factors, especially if you change more than just image size. Also when going from GLDv2c to GLDV2x, its expected to have a higher loss, as data becomes more noisy. If you only change img size, loss should converge to old level pretty fast.</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 3275451,
      "author_name": "Rhs Liza",
      "author_url": "",
      "post_date": "2025-08-26T14:54:20.903000",
      "content": "<p>Congratulations</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1561494,
      "author_name": "rahul yadav",
      "author_url": "",
      "post_date": "2021-10-27T15:26:00.907000",
      "content": "<p>Congratulations and thanks for sharing</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1542994,
      "author_name": "ysss",
      "author_url": "",
      "post_date": "2021-10-13T03:55:56.773000",
      "content": "<p>It's amazing! Thank you for sharing.</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1537908": "First, let me thank kaggle staff and google team for organizing the 2021 landmark competitions. I am still overwhelmed and shocked by my result, which resulted in becoming #1 on the overall kaggle leaderboard, an ultimate goal for many kagglers. My solutions are similar in terms of models and training routine for the retrieval and the recognition task, but slightly different in details, especially in postprocessing, i.e. in re-ranking procedure. I try to give attribution to people, whose smart ideas I levered from last year's solutions. For recognition track I re-used several ideas from 2nd place solution of GLR2020 by @bestfitting, 3rd place solution of team All Data Are Ext ( @boliu0 @haqishen @garybios @alexanderliao ) Please comment if I forget someone.\n\n## TLDR\nMy solution is an ensemble of DOLG models [1] and Hybrid-Swin-Transformers, which are both especially strong architectures for integrating local and global descriptors of a landmark into a single feature vector in a single-stage manner. I trained the models using a 3-step approach where I enlarged the training dataset and image size in each step. For postprocessing, which mainly handles non-landmark identification, I used the same method as in last year's 1st place recognition solution.\n\n## Preamble\nBefore going into details I want to address a few circumstances leading me to the solution. Firstly, compared to last year, this year's competition had a significantly more tense timeline, which forced me to skip following some ideas I had and concentrate on the most promising. That especially led to the choice to not consider spatial verification with local descriptors at all, as it has (at least in my opinion) a bad cost-benefit trade-off, especially in large scale image recognition, with limited kernel runtime. Secondly, since top teams of last year's iterations published their solutions in detail, accessible to everybody, the bar to outperform here was quite high. To be consistent in naming I use the following for subsets of gldv2\n\n**gldv2** available under [2] (5 mio images, 200k classes)\n**gldv2c** (cleaned subset of gldv2 / this years training data) (1.6 million training images, 81k classes)\n**gldv2x**, uncleaned subset of gldv2 restricted to only have the same 81k classes as gldv2c (3.2 mio images)\n\nFurthermore I use **GLRec** for past google landmark recognition competitions and **GLRet** for retrieval. \n\n## Cross-validation\nI used the same cross-validation scheme as last year, namely the 2019 labeled test set, and had a good correlation between Public LB and my local validation.\n\n## Training routine\nMy training routine follows train dataset choice from 2nd place GLRec 2020 but with some improvements using ideas from *All Data Are Ext*\nAs a first step I train models on small image size (e.g. 224x224) on the clean gldv2c for around 10 epochs. Then I use medium large image size (e.g. 512x512) and train for a long time on the more noisy gldv2x (30-40 epochs). Finally I finetune on large image size (e.g. 768x768) also on gldv2x for a few epochs. I used Adam optimizer and cosine annealing schedule in every step. I used same augmentation as *All Data Are Ext*, which proved to be very efficient:\n\n```\ncfg.train_aug = A.Compose([\n        A.HorizontalFlip(p=0.5),\n        A.ImageCompression(quality_lower=99, quality_upper=100),\n        A.ShiftScaleRotate(shift_limit=0.2, scale_limit=0.2, rotate_limit=10, border_mode=0, p=0.7),\n        A.Resize(image_size, image_size),\n        A.Cutout(max_h_size=int(image_size * 0.4), max_w_size=int(image_size * 0.4), num_holes=1, p=0.5),\n    ])\n\ncfg.val_aug = A.Compose([A.Resize(image_size, image_size),])\n\n```\n\n## Model Architectures\n\nFor me this is the most fun part in every competition, to design and explore model architectures. I like the idea of models solving multiple parts of the problem at hand in end2end fashion. In landmark recognition it is crucial to leverage both, global and local image information. While historically this was solved in a two-stage fashion by training a global descriptor for preliminary recognition on the one hand and extracting local descriptors for post reranking using spatial verification on the other hand. Recently attempts have been made to train local and global descriptors simultaneously using a single model (DELF/ DELG). But that still uses a 2nd stage spatial verification. \nI implemented two architectures that combine local and global descriptor into a single image embedding. Both use an effientnet encoder and an sub-arcface head with dynamic margins which was shown by team *All Data Are Ext* in 2020 to outperform classic arcface.\n\n### DOLG\n\nThe very recently released DOLG paper, goes on step further claiming to train local and global descriptors also simultaneously but additionally fusing them within the model into a single descriptor. Unfortunately/ fortunately it was so new that no code was released so far and I needed to implement deep orthogonal fusion by myself. Luckily the paper gives enough details and I was able to implement it quickly. The following shows the final architecture:\n\n![DOLG-EffiencentNet](https://i.imgur.com/9FiSEet.png)\n\n### Hybrid Swin Transformer\n\nAnother architecture I explored was a Hybrid-Transformer, half a CNN encoder with a transformer ending. The idea behind is that transformers are like graph neural nets on image patches connecting distant local features with each other. A property CNNs are lacking, but which are very important for landmark recognition/ retrieval. However, pretrained image transformers are mostly done on 224 or 384 image size, and I wanted to go bigger. Also the encoding of each patch is of a higher quality when using a sophisticated CNN for the patch embedding. Luckily the awesome timm repository gives you the tool set to glue together a CNN and an image transformer and I ended with an architecture like shown below. In particular I used Swin transformer [3], as I think the moving window approach captures more diverse local descriptors and EfficientNet for encoding.\n\n![Hybrid-Swin-Transformer](https://i.imgur.com/2iXNuBA.png)\n\nHowever, for training several problems arise: Firstly the pretrained weights of image transformer and CNN are trained separately, hence not fitting very well together, resulting in NaNs while training. Secondly, since image transformer require a fixed input size it's a bit more difficult to follow a training routine with increasing image size in each step. I found that the following procedure fixes the issues and gives a good final model, and especially step 3 is important to make the Hybrid-Swin-Transformer working\n\n1. train only image transformer on 224x224 size\n2. Exchange original patch embedding module with block 0,1,2 from Effnet encoder, \n3. freeze image transformer and sub-arcface head and train for 1 epoch on 448x448, to let the effnet encoder “adjust”\n4. unfreeze image transformer and head and train for 30-40 epochs\n5. add block 3 from effnet encoder and finetune a few epochs on 896x896\n\nFor two of my models I used the “stride trick”, i.e. putting a stride of (1,1) in the first conv layer to effectively double the used resolution instead of double the actual image size, which I think is better for small original images.\nI also retrained a few models from 2nd and 3rd place solutions from 2020 recognition competition following their provided git repositories and instructions, and evaluated possible ensembles on my local cross validation. Models form All Data Ext team where quite low correlated to mine and added a nice diversification benefit\n\nIn my final ensemble I had the following models:\n- DOLG-Efficentnet b5 (768x768)\n- DOLG-Efficentnet b6 (768x768)\n- DOLG-Efficentnet b7 (448x448) stride 1\n- Hybrid SwinBase224-Efficentnet b5 (448x448) stride 1\n- Hybrid SwinBase224-Efficentnet b3 (896x896)\n- Hybrid SwinBase384-Efficentnet b6 (384,384) stride 1\n- All Data Ext 2020 Efficientnet b6 (512x512) \n- All Data Ext 2020 EfficientNet b3 (768x768)\n\nPost-processing\nI literally used the same approach (developed by @philippsinger) as in our last year solution. you can find it here:\n\nhttps://www.kaggle.com/c/landmark-recognition-2020/discussion/187821\n\nFor more details please refer to the paper on arxiv. Thank you for reading.\n\n[1] https://arxiv.org/abs/2108.02927 (DOLG)\n[2] https://github.com/cvdfoundation/google-landmark\n[3] https://arxiv.org/abs/2103.14030 (Swin)\n\nCode: https://github.com/ChristofHenkel/kaggle-landmark-2021-1st-place (in progress) \nPaper: submitted to arxiv, waiting for acception. In the meanwhile I uploaded to https://github.com/ChristofHenkel/kaggle-landmark-2021-1st-place/blob/main/GLR_2021_1st_place.pdf\n\n",
    "1682949": "congrats @christofhenkel 🤩",
    "1539952": "I'm glad to know your solution. it gave me great inspiration.👍👍\nThank you😃",
    "1539070": "Congratulations @christofhenkel on winning the comp and also becoming 1st ranked GM.\nTraining these big models with so much big data is not that easy, so what was your experiment setup or your training recipe. I mean did you keep experimenting with different architecture, hyperparameters manually or were you using some tool like optuna or something to select the right hyperparameters and all.",
    "1542606": "Awesome models! Congrats on the double win. Thanks for sharing your models and solution summary.",
    "1540145": "Much congratulations @christofhenkel for this wonderful achievement!\n\nIn your paper, you mentioned that you trained your models on 8xV100 NVIDIA GPUs with distributed data parallel; did you use a cloud provider, or did you have all the GPUs in hand? For new Kagglers, what GPU(s) would you recommend us getting?\n\nAlso, how long did it take to train your entire ensemble, with distributed data parallel?",
    "1672170": "How did u label previous test sets for your validation set?",
    "1656359": "What loss function have u used?\n",
    "1585935": "Hi @christofhenkel, thank you for sharing and your models are really fascinating, especially the hybrid model, and I am trying to replicate it. However, there are some points that are a bit unclear to me, I hope you could take a bit of time to explain them. Let's focus on the \"Hybrid SwinBase224-Efficentnet b3 (896x896)\" model, I think the others are similar.\n\n1. In your write-up, you said \"Exchange original patch embedding module with block 0,1,2 from Effnet encoder\",  I think it should be \"blocks 0, 1\" only because the output of it has the size of (BS, 32, 224, 224), isn't it? Accordingly, in step 5, \"add block 3 from effnet encoder and finetune a few epochs on 896x896\", block 2 should be added, not 3, if blocks 0, 1 are used.\n\n2. As I experimented, the output size of the first 2 blocks of EfficientNet-b5 is (BS, 32, 224, 224), but the input size of the Swin Transformer should be (BS, 3136, 128), this is the output size after the \"patch_embed\" layer. \n```\n(patch_embed): PatchEmbed(\n    (proj): Conv2d(32, 128, kernel_size=(4, 4), stride=(4, 4))\n    (norm): LayerNorm((128,), eps=1e-05, elementwise_affine=True)\n  )\n```\nYou mentioned that the output of the 2-block of EfficientNet-b5 is flattened, added with positional embeddings, and projected. As far as I understand, \"flattened\" here is with the patch (4 x 4) level (not pixel level (1 x 1)), so we can't use nn.Flatten(), am I correct? After flattening and transposing with patches, the input of the Swin transformer should be (BS, 3136, 32), and now you projected it into the (BS, 3136, 128) by nn.Linear(32, 128) to fit the input dim of the transformer. If all of this is correctly understood from my side, one question is, why don't you simply input the output of the 2-block EfficientNet-B5 into the Swin transformer at the beginning (before the \"patch_embed\" layer)?\n\n3. You mentioned that you use ArcFace head with a \"dynamic\" margin, can I ask what \"dynamic\" here mean? Is it changed over training steps, epochs, or something else?\n\n4. The last question is the \"virtual_token\", I am curious how to add and initialize it into the embeddings, is it simply like adding one 128-dim vector into the (BS, 3136, 128) to become (BS, 3137, 128)?\n\nMy questions have been super long and I hope it doesn't bother you.\n\nThank you!",
    "1572295": "Congrats @christofhenkel for the double win, and also thanks for sharing your ideas.\nI have some questions about the DOLF Efficientnet. In the original paper of DOLG, the Local Branch consists of Multi-Atrous layers and a pooling layer and is followed by a self-attention module after concatenation. But in your design, do you drop the pooling layer and resample the size to half of its original feature map's size?",
    "1558204": "Congrats! I learned a lot from your solution.",
    "1557241": "@christofhenkel Congratulations on winning both landmark competitions and becoming #1 in leaderboard!\n\nIf you don't mind me asking, when is the code for your solution going to be available (it's currently in progress)? I understand that you're probably busy doing clean-up, but I'm just wondering when you think it will be ready to be uploaded. Thank you!",
    "1551516": "Wonderful!",
    "1551069": "Excellent!",
    "1550799": "It's amazing!",
    "1549571": "Congratulations @christofhenkel !!!.👍👍👍👍👍👍👍\nI have been wondering how you choose LR schedulers and optimizers in such a large data set (assuming you have enough time ), and how you can quickly verify your ideas, and whether you would consider using a small number of data sets for validation first ? \nOr could you recommend some papers so that I can find the answer by myself.  ",
    "1543113": "Nice work!",
    "1543068": "Congratulation on the win. Also this is an work",
    "1542269": "Amazing work. Congratulations @christofhenkel.",
    "1540422": "@christofhenkel Congrats on your win and also for ranking on on the Competition Pillar!  Thank you for your write up, I found it very useful how you approach a problem with a tight timeline - leveraging previous work (and thanks to for those links).  ",
    "1539996": "Congratulations!. Thanks a lot for sharing.",
    "1538043": "Congratulations @christofhenkel!\nI also tried progressive training but the loss increased for the first few epochs. Wanted to know if this is expected behaviour or something wrong with the implementation?",
    "3275451": "Congratulations",
    "1561494": "Congratulations and thanks for sharing",
    "1542994": "It's amazing! Thank you for sharing."
  }
}