{
  "id": 275876,
  "title": "5th Place Solution Sharing: A Learnable Pooling Approach",
  "url": "/competitions/landmark-recognition-2021/writeups/notenoughfitting-5th-place-solution-sharing-a-lear",
  "author_name": "",
  "post_date": "2021-10-02T03:34:23.637Z",
  "votes": 33,
  "comment_count": 8,
  "views": 0,
  "content": "<p>Congratulations to the winners and thanks google and kaggle for hosting such an interesting competition! </p>\n<p>This is my first time to participate in the landmark recognition/retrieval competition. After reading the previous winning solutions, I also felt the anxiety with the limited computation resource I have: 1x3090 at the very beginning. So I decided to find a better(affordable) model architecture for the task, instead of following the common network+gem+arc_face approach.</p>\n<p>Based on the observation that larger input images would lead to better accuracy, I tried to use the intermediate feature maps from the backbones. With my experience in the Yourtube8M video classification, I then find that NeXtVLAD(<a href=\"https://arxiv.org/pdf/1811.05014.pdf\" target=\"_blank\">https://arxiv.org/pdf/1811.05014.pdf</a>) perform pretty good in aggregating and merging features from these different feature maps. </p>\n<h3>Model Architecture</h3>\n<p><img src=\"https://i.postimg.cc/SRwWXn7s/swin-nextvlad-drawio-1.png\"></p>\n<p>As you can see, each feature map is aggregated to one high-dimensional 1D feature with NeXtVLAD, then those features are concatenated and fed into a simple linear projection layer with output_dim=512. In my implementation, I also add more non-linearity with another SE-gating layer(which was actually found to be not really helpful). At last, we use the arc face product and dynamic margin for the loss.</p>\n<p>Training setup:<br>\nV2X: 3.2M images with 81313 landmark<br>\nV2C: 1.6M images with 81313 landmark<br>\nV2Full: 4.1M images with 203094 landmarks+ nonlandmark</p>\n<p>1) fix the weight in backbone and only train the learnable pooling layer with V2X for 10 epochs, which surprisingly have very high local validation(around 90%). [AdamW, 0.001 initial, 0.8 decay)</p>\n<p>2) finetune the whole model with V2C for 10 epochs(around 94% local validation accuracy)[AdamW, 0.0001 initial, 0.9 decay)</p>\n<p>3) change only the arc_face_product layer with 203095 classes(use the nonlandmark from 2019 testset is really helpful). Again, fix the weight of backbone and only tune the learnable pooling layers with V2Full. (around 95-96% accuracy after 10 epochs in local validation)[AdamW, 0.001 initial, 0.8 decay]</p>\n<p>The advantage of the learnable pooling layer is:  </p>\n<ol>\n<li>The training is super fast as you only need to train the learnable pooling layers at step 1 and step 3.</li>\n<li>I use the image input size of 256x256 at step 1 and step 2. Then increase the image size only at step 3, when the backbone is fixed(so I never need to train an efficientnet with large input images).</li>\n</ol>\n<p>For instance, training a swin-base-224 only takes around 3 days using single 3090 GPU with the setup. </p>\n<p>I achieved top10 around early Sep with just one 3090. And then my teammate <a href=\"https://www.kaggle.com/marbury\" target=\"_blank\">@marbury</a> provided me with another 2xV100 so that I can just scale to different backbones. My final solution is an ensemble of</p>\n<ul>\n<li>swin-base-224</li>\n<li>swin-base-384</li>\n<li>swin-large-224</li>\n<li>resnest101 (512 input at the step 3)</li>\n<li>efficientv2-m (512 input at step 3)</li>\n</ul>\n<p>The final LB score is highly correlated with the accuracy of the backbone models in imagenet, which I suppose show me that this learnable pooling approach have good ability to do the transfer learning. </p>\n<p>It would be interesting to see what the model arch can achieve with more complex backbones, including efficientb5-b7, efficientv2-large and swin-large-384.  There is still plenty of room to improve and hopefully to inspire more ideas &amp; works using the learnable pooling module for the competition in next year. Will share my code once it is cleaned up.</p>\n<h3>Other Insights</h3>\n<ul>\n<li>using the whole v2full as the index set during the inference is important to stay at top 10.</li>\n</ul>",
  "messages": [
    {
      "id": "1531430",
      "postDate": "10/02/2021 02:13:58",
      "content": "<p>Congratulations to the winners and thanks google and kaggle for hosting such an interesting competition! </p>\n<p>This is my first time to participate in the landmark recognition/retrieval competition. After reading the previous winning solutions, I also felt the anxiety with the limited computation resource I have: 1x3090 at the very beginning. So I decided to find a better(affordable) model architecture for the task, instead of following the common network+gem+arc_face approach.</p>\n<p>Based on the observation that larger input images would lead to better accuracy, I tried to use the intermediate feature maps from the backbones. With my experience in the Yourtube8M video classification, I then find that NeXtVLAD(<a href=\"https://arxiv.org/pdf/1811.05014.pdf\" target=\"_blank\">https://arxiv.org/pdf/1811.05014.pdf</a>) perform pretty good in aggregating and merging features from these different feature maps. </p>\n<h3>Model Architecture</h3>\n<p><img src=\"https://i.postimg.cc/SRwWXn7s/swin-nextvlad-drawio-1.png\"></p>\n<p>As you can see, each feature map is aggregated to one high-dimensional 1D feature with NeXtVLAD, then those features are concatenated and fed into a simple linear projection layer with output_dim=512. In my implementation, I also add more non-linearity with another SE-gating layer(which was actually found to be not really helpful). At last, we use the arc face product and dynamic margin for the loss.</p>\n<p>Training setup:<br>\nV2X: 3.2M images with 81313 landmark<br>\nV2C: 1.6M images with 81313 landmark<br>\nV2Full: 4.1M images with 203094 landmarks+ nonlandmark</p>\n<p>1) fix the weight in backbone and only train the learnable pooling layer with V2X for 10 epochs, which surprisingly have very high local validation(around 90%). [AdamW, 0.001 initial, 0.8 decay)</p>\n<p>2) finetune the whole model with V2C for 10 epochs(around 94% local validation accuracy)[AdamW, 0.0001 initial, 0.9 decay)</p>\n<p>3) change only the arc_face_product layer with 203095 classes(use the nonlandmark from 2019 testset is really helpful). Again, fix the weight of backbone and only tune the learnable pooling layers with V2Full. (around 95-96% accuracy after 10 epochs in local validation)[AdamW, 0.001 initial, 0.8 decay]</p>\n<p>The advantage of the learnable pooling layer is:  </p>\n<ol>\n<li>The training is super fast as you only need to train the learnable pooling layers at step 1 and step 3.</li>\n<li>I use the image input size of 256x256 at step 1 and step 2. Then increase the image size only at step 3, when the backbone is fixed(so I never need to train an efficientnet with large input images).</li>\n</ol>\n<p>For instance, training a swin-base-224 only takes around 3 days using single 3090 GPU with the setup. </p>\n<p>I achieved top10 around early Sep with just one 3090. And then my teammate <a href=\"https://www.kaggle.com/marbury\" target=\"_blank\">@marbury</a> provided me with another 2xV100 so that I can just scale to different backbones. My final solution is an ensemble of</p>\n<ul>\n<li>swin-base-224</li>\n<li>swin-base-384</li>\n<li>swin-large-224</li>\n<li>resnest101 (512 input at the step 3)</li>\n<li>efficientv2-m (512 input at step 3)</li>\n</ul>\n<p>The final LB score is highly correlated with the accuracy of the backbone models in imagenet, which I suppose show me that this learnable pooling approach have good ability to do the transfer learning. </p>\n<p>It would be interesting to see what the model arch can achieve with more complex backbones, including efficientb5-b7, efficientv2-large and swin-large-384.  There is still plenty of room to improve and hopefully to inspire more ideas &amp; works using the learnable pooling module for the competition in next year. Will share my code once it is cleaned up.</p>\n<h3>Other Insights</h3>\n<ul>\n<li>using the whole v2full as the index set during the inference is important to stay at top 10.</li>\n</ul>",
      "rawMarkdown": "Congratulations to the winners and thanks google and kaggle for hosting such an interesting competition! \n\nThis is my first time to participate in the landmark recognition/retrieval competition. After reading the previous winning solutions, I also felt the anxiety with the limited computation resource I have: 1x3090 at the very beginning. So I decided to find a better(affordable) model architecture for the task, instead of following the common network+gem+arc_face approach.\n\nBased on the observation that larger input images would lead to better accuracy, I tried to use the intermediate feature maps from the backbones. With my experience in the Yourtube8M video classification, I then find that NeXtVLAD(https://arxiv.org/pdf/1811.05014.pdf) perform pretty good in aggregating and merging features from these different feature maps. \n\n\n### Model Architecture \n\n<img src=\"https://i.postimg.cc/SRwWXn7s/swin-nextvlad-drawio-1.png\"\nwidth=\"3000px\">\n\nAs you can see, each feature map is aggregated to one high-dimensional 1D feature with NeXtVLAD, then those features are concatenated and fed into a simple linear projection layer with output_dim=512. In my implementation, I also add more non-linearity with another SE-gating layer(which was actually found to be not really helpful). At last, we use the arc face product and dynamic margin for the loss.\n\nTraining setup:\nV2X: 3.2M images with 81313 landmark\nV2C: 1.6M images with 81313 landmark\nV2Full: 4.1M images with 203094 landmarks+ nonlandmark\n \n1) fix the weight in backbone and only train the learnable pooling layer with V2X for 10 epochs, which surprisingly have very high local validation(around 90%). [AdamW, 0.001 initial, 0.8 decay)\n\n2) finetune the whole model with V2C for 10 epochs(around 94% local validation accuracy)[AdamW, 0.0001 initial, 0.9 decay)\n\n3) change only the arc_face_product layer with 203095 classes(use the nonlandmark from 2019 testset is really helpful). Again, fix the weight of backbone and only tune the learnable pooling layers with V2Full. (around 95-96% accuracy after 10 epochs in local validation)[AdamW, 0.001 initial, 0.8 decay]\n\nThe advantage of the learnable pooling layer is:  \n\n1. The training is super fast as you only need to train the learnable pooling layers at step 1 and step 3.\n2. I use the image input size of 256x256 at step 1 and step 2. Then increase the image size only at step 3, when the backbone is fixed(so I never need to train an efficientnet with large input images).\n\nFor instance, training a swin-base-224 only takes around 3 days using single 3090 GPU with the setup. \n\nI achieved top10 around early Sep with just one 3090. And then my teammate @marbury provided me with another 2xV100 so that I can just scale to different backbones. My final solution is an ensemble of\n- swin-base-224\n- swin-base-384\n- swin-large-224\n- resnest101 (512 input at the step 3)\n- efficientv2-m (512 input at step 3)\n\nThe final LB score is highly correlated with the accuracy of the backbone models in imagenet, which I suppose show me that this learnable pooling approach have good ability to do the transfer learning. \n\nIt would be interesting to see what the model arch can achieve with more complex backbones, including efficientb5-b7, efficientv2-large and swin-large-384.  There is still plenty of room to improve and hopefully to inspire more ideas & works using the learnable pooling module for the competition in next year. Will share my code once it is cleaned up.\n\n### Other Insights\n- using the whole v2full as the index set during the inference is important to stay at top 10.",
      "votes": null
    },
    {
      "id": "1531485",
      "postDate": "10/02/2021 03:49:34",
      "content": "<p>Congratulations and thank you for sharing! 🎉</p>\n<p>I'm curious about the inference process, </p>\n<blockquote>\n  <p>using the whole v2full as the index set during the inference is important to stay at top 10.</p>\n</blockquote>\n<p>Since the private training set is a 100k subset of the V2C, will this result in a recognition result outside the 81313 classes?</p>",
      "rawMarkdown": "Congratulations and thank you for sharing! 🎉\n\nI'm curious about the inference process, \n> using the whole v2full as the index set during the inference is important to stay at top 10.\n\nSince the private training set is a 100k subset of the V2C, will this result in a recognition result outside the 81313 classes?",
      "votes": null
    },
    {
      "id": "1531491",
      "postDate": "10/02/2021 03:58:17",
      "content": "<p><a href=\"https://www.kaggle.com/apretrue\" target=\"_blank\">@apretrue</a> the test set is not a subset of V2C but V2Full. I encountered the same problem at the beginning(you can refer to my previous post in this competition).</p>\n<p>To be more precise, using a subset of the V2full during inference is important. During the inference, I upload the pre-calculated embeddings for V2Full and simply filter out the 'invalid' landmarks which didn't appear in the private index(training) set. </p>",
      "rawMarkdown": "apretrue the test set is not a subset of V2C but V2Full. I encountered the same problem at the beginning(you can refer to my previous post in this competition).\n\nTo be more precise, using a subset of the V2full during inference is important. During the inference, I upload the pre-calculated embeddings for V2Full and simply filter out the 'invalid' landmarks which didn't appear in the private index(training) set.",
      "votes": null
    },
    {
      "id": "1531548",
      "postDate": "10/02/2021 05:52:00",
      "content": "<p>Congrats on your amazingly efficient solution! Any details about inference or is it just retrieval based top1? </p>",
      "rawMarkdown": "Congrats on your amazingly efficient solution! Any details about inference or is it just retrieval based top1?",
      "votes": null
    },
    {
      "id": "1531638",
      "postDate": "10/02/2021 07:56:15",
      "content": "<p>A very interesting idea with the merge of models before the final embedding vector. Thanks for the write up</p>",
      "rawMarkdown": "A very interesting idea with the merge of models before the final embedding vector. Thanks for the write up",
      "votes": null
    },
    {
      "id": "1531840",
      "postDate": "10/02/2021 12:32:10",
      "content": "<p><a href=\"https://www.kaggle.com/keremt\" target=\"_blank\">@keremt</a>  </p>\n<p>(not sure why the reply is not working)</p>\n<p>mainly use the same inference process from last year's winning solution. Find the 3 nearest neighbors from filtered V2Full pretrained embedding and penalize by the similarity to nonlandmarks(A_ij - B_j as mentioned in the post). Finally downrank the nonlandmarks with scores &gt; 0.3. Nothing really new and interesting. I am sure you can learn more from other top solutions about the postprocessing part.</p>",
      "rawMarkdown": "keremt  \n\n(not sure why the reply is not working)\n\nmainly use the same inference process from last year's winning solution. Find the 3 nearest neighbors from filtered V2Full pretrained embedding and penalize by the similarity to nonlandmarks(A_ij - B_j as mentioned in the post). Finally downrank the nonlandmarks with scores > 0.3. Nothing really new and interesting. I am sure you can learn more from other top solutions about the postprocessing part.",
      "votes": null
    },
    {
      "id": "1531952",
      "postDate": "10/02/2021 13:43:03",
      "content": "<p>Hi congratz on the win and sharing your elegant solution here!</p>\n<p>I would like to also clarify something about the private train and test sets here.</p>\n<p>From the description in the competition <a href=\"https://www.kaggle.com/c/landmark-recognition-2021/data\" target=\"_blank\">page</a>:<br>\n<code>the private training set contains only a 100k subset of the total public training set. This 100k subset contains all of the training set images associated with the landmarks in the private test set</code></p>\n<p>But in reality, the 100k images in the private training set is <strong>not a subset</strong> of the total public training set and both the private train and private test sets contain images with landmark ids that are outside of the  81313 landmark ids present in public train. May I ask if my understanding is correct and just curious on how did you find out about this?</p>",
      "rawMarkdown": "Hi congratz on the win and sharing your elegant solution here!\n\nI would like to also clarify something about the private train and test sets here.\n\nFrom the description in the competition [page](https://www.kaggle.com/c/landmark-recognition-2021/data):\n```the private training set contains only a 100k subset of the total public training set. This 100k subset contains all of the training set images associated with the landmarks in the private test set```\n\nBut in reality, the 100k images in the private training set is **not a subset** of the total public training set and both the private train and private test sets contain images with landmark ids that are outside of the  81313 landmark ids present in public train. May I ask if my understanding is correct and just curious on how did you find out about this?",
      "votes": null
    },
    {
      "id": "1531960",
      "postDate": "10/02/2021 13:48:37",
      "content": "<p><a href=\"https://www.kaggle.com/tmxxuan\" target=\"_blank\">@tmxxuan</a>  Yes, the 100k images in private training set should come from the 200k landmarks instead of 81313 ones. I got notebook exception threw error when applied a landmark_id mask to only 81313 classes. While debugging the issue, I think \"the total public training set\" actually means the V2Full. The description is unfortunately not very clear. </p>",
      "rawMarkdown": "tmxxuan  Yes, the 100k images in private training set should come from the 200k landmarks instead of 81313 ones. I got notebook exception threw error when applied a landmark_id mask to only 81313 classes. While debugging the issue, I think \"the total public training set\" actually means the V2Full. The description is unfortunately not very clear.",
      "votes": null
    },
    {
      "id": "1531984",
      "postDate": "10/02/2021 14:08:12",
      "content": "<p>I agree that the description was not very clear. I had so many <code>Notebook Threw Exception</code> too. Fortunately, some people shared their investigation about the private dataset, for example, <a href=\"https://www.kaggle.com/c/landmark-recognition-2021/discussion/270020#1501830\" target=\"_blank\">here</a> and <a href=\"https://www.kaggle.com/c/landmark-recognition-2021/discussion/273849#1521980\" target=\"_blank\">here</a>. Such discussion helped me a lot to improve the public score. </p>",
      "rawMarkdown": "I agree that the description was not very clear. I had so many `Notebook Threw Exception` too. Fortunately, some people shared their investigation about the private dataset, for example, [here](https://www.kaggle.com/c/landmark-recognition-2021/discussion/270020#1501830) and [here](https://www.kaggle.com/c/landmark-recognition-2021/discussion/273849#1521980). Such discussion helped me a lot to improve the public score.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1531485,
      "author_name": "apretrue",
      "author_url": "",
      "post_date": "10/02/2021 03:49:34",
      "content": "<p>Congratulations and thank you for sharing! 🎉</p>\n<p>I'm curious about the inference process, </p>\n<blockquote>\n  <p>using the whole v2full as the index set during the inference is important to stay at top 10.</p>\n</blockquote>\n<p>Since the private training set is a 100k subset of the V2C, will this result in a recognition result outside the 81313 classes?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1531491,
          "author_name": "phoenixl",
          "author_url": "",
          "post_date": "10/02/2021 03:58:17",
          "content": "<p><a href=\"https://www.kaggle.com/apretrue\" target=\"_blank\">@apretrue</a> the test set is not a subset of V2C but V2Full. I encountered the same problem at the beginning(you can refer to my previous post in this competition).</p>\n<p>To be more precise, using a subset of the V2full during inference is important. During the inference, I upload the pre-calculated embeddings for V2Full and simply filter out the 'invalid' landmarks which didn't appear in the private index(training) set. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1531952,
          "author_name": "tmxxuan",
          "author_url": "",
          "post_date": "10/02/2021 13:43:03",
          "content": "<p>Hi congratz on the win and sharing your elegant solution here!</p>\n<p>I would like to also clarify something about the private train and test sets here.</p>\n<p>From the description in the competition <a href=\"https://www.kaggle.com/c/landmark-recognition-2021/data\" target=\"_blank\">page</a>:<br>\n<code>the private training set contains only a 100k subset of the total public training set. This 100k subset contains all of the training set images associated with the landmarks in the private test set</code></p>\n<p>But in reality, the 100k images in the private training set is <strong>not a subset</strong> of the total public training set and both the private train and private test sets contain images with landmark ids that are outside of the  81313 landmark ids present in public train. May I ask if my understanding is correct and just curious on how did you find out about this?</p>",
          "votes": null,
          "replies": [
            {
              "id": 1531960,
              "author_name": "phoenixl",
              "author_url": "",
              "post_date": "10/02/2021 13:48:37",
              "content": "<p><a href=\"https://www.kaggle.com/tmxxuan\" target=\"_blank\">@tmxxuan</a>  Yes, the 100k images in private training set should come from the 200k landmarks instead of 81313 ones. I got notebook exception threw error when applied a landmark_id mask to only 81313 classes. While debugging the issue, I think \"the total public training set\" actually means the V2Full. The description is unfortunately not very clear. </p>",
              "votes": null,
              "replies": []
            }
          ]
        },
        {
          "id": 1531984,
          "author_name": "hdsk38",
          "author_url": "",
          "post_date": "10/02/2021 14:08:12",
          "content": "<p>I agree that the description was not very clear. I had so many <code>Notebook Threw Exception</code> too. Fortunately, some people shared their investigation about the private dataset, for example, <a href=\"https://www.kaggle.com/c/landmark-recognition-2021/discussion/270020#1501830\" target=\"_blank\">here</a> and <a href=\"https://www.kaggle.com/c/landmark-recognition-2021/discussion/273849#1521980\" target=\"_blank\">here</a>. Such discussion helped me a lot to improve the public score. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1531548,
      "author_name": "keremt",
      "author_url": "",
      "post_date": "10/02/2021 05:52:00",
      "content": "<p>Congrats on your amazingly efficient solution! Any details about inference or is it just retrieval based top1? </p>",
      "votes": null,
      "replies": [
        {
          "id": 1531840,
          "author_name": "phoenixl",
          "author_url": "",
          "post_date": "10/02/2021 12:32:10",
          "content": "<p><a href=\"https://www.kaggle.com/keremt\" target=\"_blank\">@keremt</a>  </p>\n<p>(not sure why the reply is not working)</p>\n<p>mainly use the same inference process from last year's winning solution. Find the 3 nearest neighbors from filtered V2Full pretrained embedding and penalize by the similarity to nonlandmarks(A_ij - B_j as mentioned in the post). Finally downrank the nonlandmarks with scores &gt; 0.3. Nothing really new and interesting. I am sure you can learn more from other top solutions about the postprocessing part.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1531638,
      "author_name": "narsil",
      "author_url": "",
      "post_date": "10/02/2021 07:56:15",
      "content": "<p>A very interesting idea with the merge of models before the final embedding vector. Thanks for the write up</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1531430": "Congratulations to the winners and thanks google and kaggle for hosting such an interesting competition! \n\nThis is my first time to participate in the landmark recognition/retrieval competition. After reading the previous winning solutions, I also felt the anxiety with the limited computation resource I have: 1x3090 at the very beginning. So I decided to find a better(affordable) model architecture for the task, instead of following the common network+gem+arc_face approach.\n\nBased on the observation that larger input images would lead to better accuracy, I tried to use the intermediate feature maps from the backbones. With my experience in the Yourtube8M video classification, I then find that NeXtVLAD(https://arxiv.org/pdf/1811.05014.pdf) perform pretty good in aggregating and merging features from these different feature maps. \n\n\n### Model Architecture \n\n<img src=\"https://i.postimg.cc/SRwWXn7s/swin-nextvlad-drawio-1.png\"\nwidth=\"3000px\">\n\nAs you can see, each feature map is aggregated to one high-dimensional 1D feature with NeXtVLAD, then those features are concatenated and fed into a simple linear projection layer with output_dim=512. In my implementation, I also add more non-linearity with another SE-gating layer(which was actually found to be not really helpful). At last, we use the arc face product and dynamic margin for the loss.\n\nTraining setup:\nV2X: 3.2M images with 81313 landmark\nV2C: 1.6M images with 81313 landmark\nV2Full: 4.1M images with 203094 landmarks+ nonlandmark\n \n1) fix the weight in backbone and only train the learnable pooling layer with V2X for 10 epochs, which surprisingly have very high local validation(around 90%). [AdamW, 0.001 initial, 0.8 decay)\n\n2) finetune the whole model with V2C for 10 epochs(around 94% local validation accuracy)[AdamW, 0.0001 initial, 0.9 decay)\n\n3) change only the arc_face_product layer with 203095 classes(use the nonlandmark from 2019 testset is really helpful). Again, fix the weight of backbone and only tune the learnable pooling layers with V2Full. (around 95-96% accuracy after 10 epochs in local validation)[AdamW, 0.001 initial, 0.8 decay]\n\nThe advantage of the learnable pooling layer is:  \n\n1. The training is super fast as you only need to train the learnable pooling layers at step 1 and step 3.\n2. I use the image input size of 256x256 at step 1 and step 2. Then increase the image size only at step 3, when the backbone is fixed(so I never need to train an efficientnet with large input images).\n\nFor instance, training a swin-base-224 only takes around 3 days using single 3090 GPU with the setup. \n\nI achieved top10 around early Sep with just one 3090. And then my teammate @marbury provided me with another 2xV100 so that I can just scale to different backbones. My final solution is an ensemble of\n- swin-base-224\n- swin-base-384\n- swin-large-224\n- resnest101 (512 input at the step 3)\n- efficientv2-m (512 input at step 3)\n\nThe final LB score is highly correlated with the accuracy of the backbone models in imagenet, which I suppose show me that this learnable pooling approach have good ability to do the transfer learning. \n\nIt would be interesting to see what the model arch can achieve with more complex backbones, including efficientb5-b7, efficientv2-large and swin-large-384.  There is still plenty of room to improve and hopefully to inspire more ideas & works using the learnable pooling module for the competition in next year. Will share my code once it is cleaned up.\n\n### Other Insights\n- using the whole v2full as the index set during the inference is important to stay at top 10.",
    "1531485": "Congratulations and thank you for sharing! 🎉\n\nI'm curious about the inference process, \n> using the whole v2full as the index set during the inference is important to stay at top 10.\n\nSince the private training set is a 100k subset of the V2C, will this result in a recognition result outside the 81313 classes?",
    "1531491": "apretrue the test set is not a subset of V2C but V2Full. I encountered the same problem at the beginning(you can refer to my previous post in this competition).\n\nTo be more precise, using a subset of the V2full during inference is important. During the inference, I upload the pre-calculated embeddings for V2Full and simply filter out the 'invalid' landmarks which didn't appear in the private index(training) set.",
    "1531548": "Congrats on your amazingly efficient solution! Any details about inference or is it just retrieval based top1?",
    "1531638": "A very interesting idea with the merge of models before the final embedding vector. Thanks for the write up",
    "1531840": "keremt  \n\n(not sure why the reply is not working)\n\nmainly use the same inference process from last year's winning solution. Find the 3 nearest neighbors from filtered V2Full pretrained embedding and penalize by the similarity to nonlandmarks(A_ij - B_j as mentioned in the post). Finally downrank the nonlandmarks with scores > 0.3. Nothing really new and interesting. I am sure you can learn more from other top solutions about the postprocessing part.",
    "1531952": "Hi congratz on the win and sharing your elegant solution here!\n\nI would like to also clarify something about the private train and test sets here.\n\nFrom the description in the competition [page](https://www.kaggle.com/c/landmark-recognition-2021/data):\n```the private training set contains only a 100k subset of the total public training set. This 100k subset contains all of the training set images associated with the landmarks in the private test set```\n\nBut in reality, the 100k images in the private training set is **not a subset** of the total public training set and both the private train and private test sets contain images with landmark ids that are outside of the  81313 landmark ids present in public train. May I ask if my understanding is correct and just curious on how did you find out about this?",
    "1531960": "tmxxuan  Yes, the 100k images in private training set should come from the 200k landmarks instead of 81313 ones. I got notebook exception threw error when applied a landmark_id mask to only 81313 classes. While debugging the issue, I think \"the total public training set\" actually means the V2Full. The description is unfortunately not very clear.",
    "1531984": "I agree that the description was not very clear. I had so many `Notebook Threw Exception` too. Fortunately, some people shared their investigation about the private dataset, for example, [here](https://www.kaggle.com/c/landmark-recognition-2021/discussion/270020#1501830) and [here](https://www.kaggle.com/c/landmark-recognition-2021/discussion/273849#1521980). Such discussion helped me a lot to improve the public score."
  },
  "source": "meta"
}